Actions delays in starting runs
Final update
On August 24, 2026, between 13:33 UTC and 14:04 UTC, 3.8% of Actions runs experienced start delays over 5 minutes with 1.25% of Actions runs failing outright. The incident was caused by a disk failure on a node hosting one of many service instances responsible for processing runner assignment events. Typically, pods on unhealthy nodes are removed and replaced automatically without impact. In this case, although the node was severely degraded and unable to perform disk operations, it continued sending healthy signals, preventing the system from immediately moving its work elsewhere. During this period, events assigned to the affected component accumulated until an automatic rebalance redirected processing to healthy components at 13:54 UTC. The queue backlog was cleared at 14:00 UTC, and processing returned to normal by 14:04 UTC. To prevent a recurrence, we are improving detection and automated remediation for unhealthy nodes that aren’t fully offline. We are also strengthening application-level resiliency, so stalled consumers are automatically removed quickly and their work reassigned without waiting for the affected node to recover.
Timeline
- Resolved · Aug 24, 14:34 UTC
On August 24, 2026, between 13:33 UTC and 14:04 UTC, 3.8% of Actions runs experienced start delays over 5 minutes with 1.25% of Actions runs failing outright. The incident was caused by a disk failure on a node hosting one of many service instances responsible for processing runner assignment events. Typically, pods on unhealthy nodes are removed and replaced automatically without impact. In this case, although the node was severely degraded and unable to perform disk operations, it continued sending healthy signals, preventing the system from immediately moving its work elsewhere. During this period, events assigned to the affected component accumulated until an automatic rebalance redirected processing to healthy components at 13:54 UTC. The queue backlog was cleared at 14:00 UTC, and processing returned to normal by 14:04 UTC. To prevent a recurrence, we are improving detection and automated remediation for unhealthy nodes that aren’t fully offline. We are also strengthening application-level resiliency, so stalled consumers are automatically removed quickly and their work reassigned without waiting for the affected node to recover.
- Monitoring · Aug 24, 14:26 UTC
The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.
- Investigating · Aug 24, 14:22 UTC
Failures while queuing and running Actions jobs for a subset of customers are now resolving. We are monitoring for full recovery.
- Investigating · Aug 24, 13:56 UTC
We are investigating reports of degraded performance for Actions
More from GitHub
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 2310:11 UTC | Incident across several services | minor | Ongoing |
| Sep 2022:13 UTC | Incident with Pull Requests | minor | 1h 9m |
| Sep 1720:59 UTC | Elevated rate of errors for OpenAI models provided by Copilot | minor | 50m |
| Sep 1607:20 UTC | Degradation with Gemini 3.8 Flash | major | 10h 28m |
| Sep 1519:11 UTC | Disruption with some GitHub services | minor | 49m |
| Sep 1509:47 UTC | Disruption with some GitHub services | minor | 1h 30m |
Also caused by storage
All| Started | Vendor | Incident | Impact | Duration |
|---|---|---|---|---|
| Sep 1715:50 UTC | AutoOps node metrics temporarily unavailable in some regions | major | 41m | |
| Sep 1015:26 UTC | Unresponsive Projects | major | 27h 40m | |
| Jul 1909:51 UTC | Block Storage Volume NYC1, NYC3, SGP1, SYD1 and BLR1 | minor | 5h 41m | |
| Mar 720:07 UTC | Outage for prod-eu-central-0 due to AWS S3 outage. | none | 36h 52m | |
| Mar 719:53 UTC | Increased Error Rates | major | 1h 11m | |
| Nov 2808:28 UTC | Delayed Logs in EU Region | minor | 10h 46m |
From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.