Internal network congestion disrupts US-EAST-1
Dec 7, 15:30 UTCDec 7, 22:40 UTC
Duration
7h 10m
Impact
Critical
Root cause
Network
AWS, 90 days
12 incidents
Affected
EC2 APIsInternal DNSConnectDynamoDBus-east-1
Lesson: Retry storms turn a small change into congestion; clients need backoff, and monitoring must not share the network it monitors.
What happened
An automated activity to scale capacity of a service in the main AWS network triggered unexpected behavior from a large number of clients inside the internal network, congesting the devices between the internal and main networks. DNS resolution recovered at 17:28 UTC and network devices at 22:22 UTC.
More from AWS
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 2123:26 UTC | Increased Error Rates | minor | 58m |
| Sep 321:49 UTC | Increased API Error Rates | minor | 2h 18m |
| Aug 2102:02 UTC | Increased Error Rates | minor | 38m |
| Aug 1915:15 UTC | Increased Error Rates | minor | 3h 32m |
| Aug 1503:42 UTC | Increased Packet loss | minor | 3d |
| Jul 3117:33 UTC | Elevated Packet Loss | minor | 1h 21m |
Also caused by network
All| Started | Vendor | Incident | Impact | Duration |
|---|---|---|---|---|
| Sep 1901:36 UTC | Elevated errors in Ashburn, VA (IAD) | none | 0m | |
| Sep 1411:59 UTC | Increased wait times for macOS jobs | minor | 1h | |
| Sep 410:57 UTC | Increased errors on High Performance Edge Network - FRA Region | minor | 0m | |
| Sep 214:18 UTC | Upstream network issues | minor | 1h 39m | |
| Sep 116:11 UTC | Service degradation in GCP us-central-1 | none | 19m | |
| Sep 114:44 UTC | Multiple products in us-central1-b are experiencing network service degradation. | major | 4h 8m |
From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.