Skip to content
Slack ยท Apps and SaaSJan 4, 2021, 14:57 UTC

Overloaded AWS Transit Gateway takes Slack down on the first workday of 2021

CriticalCapacityUpdated 19h ago
Jan 4, 14:57 UTCJan 4, 17:15 UTC
Duration
2h 18m
Impact
Critical
Root cause
Capacity
Slack, 90 days
7 incidents
Affected
MessagingWeb tierGlobal

Lesson: Managed network components have scaling limits too; pre-warm for known traffic spikes and make sure provisioning survives the incident it is meant to fix.

What happened

One of the AWS Transit Gateways linking Slack VPCs became overloaded as traffic returned after the holidays. Packet loss saturated the web tier, autoscaling tried to add 1,200 servers, and the provisioning service itself overloaded before the web tier recovered.

More from Slack

Full history

Also caused by capacity and load

All
StartedIncidentDuration
Sep 2213:19 UTCWe are investigating an issue with CH servers in Azure germanywestcentral regionClickHouse5h 6m
Sep 1607:20 UTCDegradation with Gemini 3.8 FlashGitHub10h 28m
Sep 1509:47 UTCDisruption with some GitHub servicesGitHub1h 30m
Sep 422:02 UTCDegradation in repos contents APIGitHub21m
Sep 221:44 UTCElevated Linux worker queue timesExpo21m
Sep 115:00 UTCDelays in commit processingGitHub1h 1m

From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.

Weekly: the week's major outages, postmortems and breaches, Saturday mornings.