Incident with Actions
Lesson: Teams should monitor service status when experiencing issues with automated workflow systems.
What happened
On August 26, 2026 from 15:02 to 15:45 UTC, Actions jobs failed to start. The following 2 hours until 17:40 UTC, Actions runs were delayed starting by more than 5 minutes as the system caught up with delayed load. This impact was triggered by saturation of writes to the database primary used by the service processing triggers for Actions workflows. The primary was failed over, but the system did not fully recover. The saturation was caused by growing daily peak load combined with an upstream issue in GitHub’s event processing infrastructure, https://www.githubstatus.com/incidents/hcbtzksccj2f, which caused burst amplification of already-high load. Downstream throttles that were later used to recover were set ~10% too high to protect the system. At 15:45 UTC, throttling combined with service restarts recovered the service’s core health. Those throttles were gradually raised between 15:54 and 17:22 to restore full webhook processing for Actions runs. This ramp was deliberately slow to ensure we did not re-overwhelm the system given our original throttling was now known to be incorrectly set. The queue of webhook events was fully burned down at 17:40 UTC. 3.7% of larger-runner jobs,
Timeline
- Resolved · Aug 26, 18:01 UTC
On August 26, 2026 from 15:02 to 15:45 UTC, Actions jobs failed to start. The following 2 hours until 17:40 UTC, Actions runs were delayed starting by more than 5 minutes as the system caught up with delayed load. This impact was triggered by saturation of writes to the database primary used by the service processing triggers for Actions workflows. The primary was failed over, but the system did not fully recover. The saturation was caused by growing daily peak load combined with an upstream issue in GitHub’s event processing infrastructure, https://www.githubstatus.com/incidents/hcbtzksccj2f, which caused burst amplification of already-high load. Downstream throttles that were later used to recover were set ~10% too high to protect the system. At 15:45 UTC, throttling combined with service restarts recovered the service’s core health. Those throttles were gradually raised between 15:54 and 17:22 to restore full webhook processing for Actions runs. This ramp was deliberately slow to ensure we did not re-overwhelm the system given our original throttling was now known to be incorrectly set. The queue of webhook events was fully burned down at 17:40 UTC. 3.7% of larger-runner jobs, along with some scale-set self-hosted jobs, remained stuck in queued or “waiting for runner” state. We deployed a change to force-revoke jobs in this state, and they transitioned to failed at 18:40 UTC, about 50 minutes after incident mitigation. Releasing these jobs also freed hosted concurrency f
- Monitoring · Aug 26, 18:00 UTC
All inbound queues have recovered and Actions is operating as expected. 3.7% of jobs assigned to larger runners during the early stage of this incident are stuck waiting for runner assignment. Those will be canceled within the hour. Other runners are successfully processing all new jobs.
- Monitoring · Aug 26, 17:54 UTC
The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.
- Investigating · Aug 26, 17:32 UTC
We are continuing to observe recovery and expect actions inbound queues to be back to normal in <30min. Work will continue to flow through the system subject to per-customer concurrency limits.
- Investigating · Aug 26, 16:50 UTC
We are continuing to observe recovery and delayed queues are burning down. Some customers will continue to see increased delays until all throttled work has been completed - we expect this within the next hour.
- Investigating · Aug 26, 16:49 UTC
Pages is operating normally.
- Investigating · Aug 26, 16:14 UTC
We believe we've identified and addressed the issue and are ramping traffic back up slowly to ensure it doesn't recur. Some customers will continue to see delays as we ramp up.
- Investigating · Aug 26, 15:48 UTC
primary failover briefly improved performance but did not fully mitigate, we've throttled inbound traffic and are investigating upstream Vitess issues
- Investigating · Aug 26, 15:23 UTC
We've identified an issue with a database primary and are failing over to a replica immediately
- Investigating · Aug 26, 15:12 UTC
Pages is experiencing degraded performance. We are continuing to investigate.
- Investigating · Aug 26, 15:11 UTC
We are investigating reports of degraded availability for Actions
More from GitHub
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 2310:11 UTC | Incident across several services | minor | Ongoing |
| Sep 2022:13 UTC | Incident with Pull Requests | minor | 1h 9m |
| Sep 1720:59 UTC | Elevated rate of errors for OpenAI models provided by Copilot | minor | 50m |
| Sep 1607:20 UTC | Degradation with Gemini 3.8 Flash | major | 10h 28m |
| Sep 1519:11 UTC | Disruption with some GitHub services | minor | 49m |
| Sep 1509:47 UTC | Disruption with some GitHub services | minor | 1h 30m |
Also caused by capacity and load
All| Started | Vendor | Incident | Impact | Duration |
|---|---|---|---|---|
| Sep 2213:19 UTC | We are investigating an issue with CH servers in Azure germanywestcentral region | minor | 5h 6m | |
| Sep 221:44 UTC | Elevated Linux worker queue times | minor | 21m | |
| Aug 2506:53 UTC | [Critical] Issue with Downloads | critical | 44m | |
| Aug 2015:40 UTC | We are investigating an issue where customers may experience timeouts, service degradations, errors, and elevated latencies across multiple products in the us-west1 region. | critical | 3h 40m | |
| Aug 1812:58 UTC | INC20000163 | major | 1h 54m | |
| Aug 1713:19 UTC | No capacity in ARN | none | 2h 22m |
From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.