Skip to content
AWS ยท Cloud and hostingDec 7, 2021, 15:30 UTC

Internal network congestion disrupts US-EAST-1

CriticalNetworkUpdated 15h ago
Dec 7, 15:30 UTCDec 7, 22:40 UTC
Duration
7h 10m
Impact
Critical
Root cause
Network
AWS, 90 days
12 incidents
Affected
EC2 APIsInternal DNSConnectDynamoDBus-east-1

Lesson: Retry storms turn a small change into congestion; clients need backoff, and monitoring must not share the network it monitors.

What happened

An automated activity to scale capacity of a service in the main AWS network triggered unexpected behavior from a large number of clients inside the internal network, congesting the devices between the internal and main networks. DNS resolution recovered at 17:28 UTC and network devices at 22:22 UTC.

More from AWS

Full history
StartedIncidentDuration
Sep 2123:26 UTCIncreased Error Rates58m
Sep 321:49 UTCIncreased API Error Rates2h 18m
Aug 2102:02 UTCIncreased Error Rates38m
Aug 1915:15 UTCIncreased Error Rates3h 32m
Aug 1503:42 UTCIncreased Packet loss3d
Jul 3117:33 UTCElevated Packet Loss1h 21m

Also caused by network

All
StartedIncidentDuration
Sep 1901:36 UTCElevated errors in Ashburn, VA (IAD)Cloudflare0m
Sep 1411:59 UTCIncreased wait times for macOS jobsCircleCI1h
Sep 410:57 UTCIncreased errors on High Performance Edge Network - FRA RegionNetlify0m
Sep 214:18 UTCUpstream network issuesFly.io1h 39m
Sep 116:11 UTCService degradation in GCP us-central-1ClickHouse19m
Sep 114:44 UTCMultiple products in us-central1-b are experiencing network service degradation.Google Cloud4h 8m

From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.

Weekly: the week's major outages, postmortems and breaches, Saturday mornings.