Increased 5xx Errors
Final update
Between 12:45 AM and 4:18 AM PDT, we experienced increased 5xx errors for CloudFront customers utilizing VPC Origins connectivity. Our engineers were automatically engaged and immediately began investigating the root cause. By 2:57 AM PDT, we identified the root cause of the issue as an internal constraint on the fleet that manages connections to private VPC origins. When this constraint was reached, the system responsible for distributing routing configuration to our network processors failed to load the updated configuration data correctly, affecting routing of VPC Origin connections. At 3:52 AM PDT, we took multiple mitigation actions that led to to full recovery at 4:18 AM PDT. Now that the issue has been mitigated, customers who temporarily changed their origin type can safely revert these changes. Customers utilizing other origin types were not affected by this issue. The issue has been resolved and the service is operating normally.
Timeline
- Resolved · Jul 16, 12:21 UTC
Between 12:45 AM and 4:18 AM PDT, we experienced increased 5xx errors for CloudFront customers utilizing VPC Origins connectivity. Our engineers were automatically engaged and immediately began investigating the root cause. By 2:57 AM PDT, we identified the root cause of the issue as an internal constraint on the fleet that manages connections to private VPC origins. When this constraint was reached, the system responsible for distributing routing configuration to our network processors failed to load the updated configuration data correctly, affecting routing of VPC Origin connections. At 3:52 AM PDT, we took multiple mitigation actions that led to to full recovery at 4:18 AM PDT. Now that the issue has been mitigated, customers who temporarily changed their origin type can safely revert these changes. Customers utilizing other origin types were not affected by this issue. The issue has been resolved and the service is operating normally.
- Update · Jul 16, 11:57 UTC
We continue to see significant signs of recovery as a result of our mitigation efforts, with full recovery expected within the next 45 minutes.
- Update · Jul 16, 11:27 UTC
We are seeing initial signs of recovery and continue to work toward full recovery.
- Update · Jul 16, 11:16 UTC
We continue working to resolve the increased 5xx errors for CloudFront customers utilizing VPC Origins connectivity. Customers utilizing other origin types remain unaffected by this issue. We have further scoped the issue down to routing table capacity within the packet processing subsystem responsible for routing requests from CloudFront's edge locations to resources within customer VPCs. We have identified and are currently testing a mitigation strategy to resolve the issue. Once testing is complete, we will deploy the mitigation in a phased approach. Based on the results from these tests, we will provide a clearer estimated time for resolution in our next update. We continue to recommend that customers who are able to do so temporarily change their origin type to resolve the errors. We will provide another update by 5:15 AM PDT, or sooner if additional information becomes available.
- Update · Jul 16, 10:18 UTC
We continue working to resolve the increased 5xx errors for CloudFront customers utilizing VPC Origins connectivity. Customers utilizing other origin types remain unaffected by this issue. Based on our investigation, we believe the root cause is related to a packet processing subsystem responsible for routing requests from CloudFront's edge locations to resources within customer VPCs. We continue to recommend that customers who are able to do so temporarily change their origin type to resolve the errors. We will provide another update by 4:15 AM PDT, or sooner if additional information becomes available.
- Update · Jul 16, 09:21 UTC
Starting at 12:45 AM PDT, we are experiencing increased 5xx errors for CloudFront customers utilizing VPC Origins connectivity. We have confirmed that customers utilizing other origin types are not impacted by this issue. Our engineers are engaged and are actively working to mitigate impact. As a workaround, customers who do not require VPC Origins can change their origin type to resolve the errors. We will provide another update by 3:15 AM PDT, or sooner if more information becomes available.
- Update · Jul 16, 08:44 UTC
We are investigating increased 5xx errors for Cloudfront customers utilizing VPC Origins connectivity.
More from AWS
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 2123:26 UTC | Increased Error Rates | minor | 58m |
| Sep 321:49 UTC | Increased API Error Rates | minor | 2h 18m |
| Aug 2102:02 UTC | Increased Error Rates | minor | 38m |
| Aug 1915:15 UTC | Increased Error Rates | minor | 3h 32m |
| Aug 1503:42 UTC | Increased Packet loss | minor | 3d |
| Jul 3117:33 UTC | Elevated Packet Loss | minor | 1h 21m |
Also caused by network
All| Started | Vendor | Incident | Impact | Duration |
|---|---|---|---|---|
| Sep 1901:36 UTC | Elevated errors in Ashburn, VA (IAD) | none | 0m | |
| Sep 1411:59 UTC | Increased wait times for macOS jobs | minor | 1h | |
| Sep 410:57 UTC | Increased errors on High Performance Edge Network - FRA Region | minor | 0m | |
| Sep 214:18 UTC | Upstream network issues | minor | 1h 39m | |
| Sep 116:11 UTC | Service degradation in GCP us-central-1 | none | 19m | |
| Sep 114:44 UTC | Multiple products in us-central1-b are experiencing network service degradation. | major | 4h 8m |
From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.