Skip to content
AWS ยท Cloud and hostingFeb 28, 2017, 17:37 UTC

Mistyped command removes S3 index servers in US-EAST-1

CriticalOperatorUpdated 17h ago
Feb 28, 17:37 UTCFeb 28, 21:54 UTC
Duration
4h 17m
Impact
Critical
Root cause
Operator
AWS, 90 days
12 incidents
Affected
S3Service Health Dashboardus-east-1

Lesson: Tools should refuse to remove capacity below a safe minimum, and the status page cannot depend on the system it reports on.

What happened

An S3 team member ran an established playbook command with one input entered incorrectly, removing far more servers than intended, including those behind the index and placement subsystems. Both needed full restarts; the status dashboard itself depended on S3.

More from AWS

Full history
StartedIncidentDuration
Sep 2123:26 UTCIncreased Error Rates58m
Sep 321:49 UTCIncreased API Error Rates2h 18m
Aug 2102:02 UTCIncreased Error Rates38m
Aug 1915:15 UTCIncreased Error Rates3h 32m
Aug 1503:42 UTCIncreased Packet loss3d
Jul 3117:33 UTCElevated Packet Loss1h 21m

Also caused by operator action

All
StartedIncidentDuration
Feb 608:14 UTCR2 object storage disabled during a phishing report remediationCloudflare1h 22m
Apr 507:38 UTCMaintenance script deletes 883 customer sitesAtlassian12d 16h
Jan 3123:00 UTCPrimary database data accidentally deleted, 18-hour restoreGitLab19h

From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.

Weekly: the week's major outages, postmortems and breaches, Saturday mornings.