Skip to content
GitLab ยท Developer toolsJan 31, 2017, 23:00 UTC

Primary database data accidentally deleted, 18-hour restore

CriticalOperatorUpdated 18h ago
Jan 31, 23:00 UTCFeb 1, 18:00 UTC
Duration
19h
Impact
Critical
Root cause
Operator
GitLab, 90 days
0 incidents
Affected
GitLab.comGlobal

Lesson: A backup is only real once a restore has been tested; destructive commands on production hosts need an unmistakable prompt.

What happened

While fighting replication lag under spam load, an engineer accidentally removed data from the primary database server. Backups had not been working, so GitLab.com was restored from a six-hour-old staging snapshot on slower disks, losing projects, issues and comments made in that window.

Also caused by operator action

All
StartedIncidentDuration
Feb 608:14 UTCR2 object storage disabled during a phishing report remediationCloudflare1h 22m
Apr 507:38 UTCMaintenance script deletes 883 customer sitesAtlassian12d 16h
Feb 2817:37 UTCMistyped command removes S3 index servers in US-EAST-1AWS4h 17m

From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.

Weekly: the week's major outages, postmortems and breaches, Saturday mornings.