Skip to content
GitHub · Developer toolsAug 26, 2026, 15:11 UTC

Incident with Actions

CriticalCapacityUpdated 5h ago
Aug 26, 15:11 UTCAug 26, 18:01 UTC
Duration
2h 50m
Impact
Critical
Root cause
Capacity
GitHub, 90 days
68 incidents
Affected
ActionsPages

Lesson: Teams should monitor service status when experiencing issues with automated workflow systems.

What happened

On August 26, 2026 from 15:02 to 15:45 UTC, Actions jobs failed to start. The following 2 hours until 17:40 UTC, Actions runs were delayed starting by more than 5 minutes as the system caught up with delayed load. This impact was triggered by saturation of writes to the database primary used by the service processing triggers for Actions workflows. The primary was failed over, but the system did not fully recover. The saturation was caused by growing daily peak load combined with an upstream issue in GitHub’s event processing infrastructure, https://www.githubstatus.com/incidents/hcbtzksccj2f, which caused burst amplification of already-high load. Downstream throttles that were later used to recover were set ~10% too high to protect the system. At 15:45 UTC, throttling combined with service restarts recovered the service’s core health. Those throttles were gradually raised between 15:54 and 17:22 to restore full webhook processing for Actions runs. This ramp was deliberately slow to ensure we did not re-overwhelm the system given our original throttling was now known to be incorrectly set. The queue of webhook events was fully burned down at 17:40 UTC. 3.7% of larger-runner jobs,

Timeline

  1. Resolved · Aug 26, 18:01 UTC
    On August 26, 2026 from 15:02 to 15:45 UTC, Actions jobs failed to start. The following 2 hours until 17:40 UTC, Actions runs were delayed starting by more than 5 minutes as the system caught up with delayed load. This impact was triggered by saturation of writes to the database primary used by the service processing triggers for Actions workflows. The primary was failed over, but the system did not fully recover. The saturation was caused by growing daily peak load combined with an upstream issue in GitHub’s event processing infrastructure, https://www.githubstatus.com/incidents/hcbtzksccj2f, which caused burst amplification of already-high load. Downstream throttles that were later used to recover were set ~10% too high to protect the system. At 15:45 UTC, throttling combined with service restarts recovered the service’s core health. Those throttles were gradually raised between 15:54 and 17:22 to restore full webhook processing for Actions runs. This ramp was deliberately slow to ensure we did not re-overwhelm the system given our original throttling was now known to be incorrectly set. The queue of webhook events was fully burned down at 17:40 UTC. 3.7% of larger-runner jobs, along with some scale-set self-hosted jobs, remained stuck in queued or “waiting for runner” state. We deployed a change to force-revoke jobs in this state, and they transitioned to failed at 18:40 UTC, about 50 minutes after incident mitigation. Releasing these jobs also freed hosted concurrency f
  2. Monitoring · Aug 26, 18:00 UTC
    All inbound queues have recovered and Actions is operating as expected. 3.7% of jobs assigned to larger runners during the early stage of this incident are stuck waiting for runner assignment. Those will be canceled within the hour. Other runners are successfully processing all new jobs.
  3. Monitoring · Aug 26, 17:54 UTC
    The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.
  4. Investigating · Aug 26, 17:32 UTC
    We are continuing to observe recovery and expect actions inbound queues to be back to normal in <30min. Work will continue to flow through the system subject to per-customer concurrency limits.
  5. Investigating · Aug 26, 16:50 UTC
    We are continuing to observe recovery and delayed queues are burning down. Some customers will continue to see increased delays until all throttled work has been completed - we expect this within the next hour.
  6. Investigating · Aug 26, 16:49 UTC
    Pages is operating normally.
  7. Investigating · Aug 26, 16:14 UTC
    We believe we've identified and addressed the issue and are ramping traffic back up slowly to ensure it doesn't recur. Some customers will continue to see delays as we ramp up.
  8. Investigating · Aug 26, 15:48 UTC
    primary failover briefly improved performance but did not fully mitigate, we've throttled inbound traffic and are investigating upstream Vitess issues
  9. Investigating · Aug 26, 15:23 UTC
    We've identified an issue with a database primary and are failing over to a replica immediately
  10. Investigating · Aug 26, 15:12 UTC
    Pages is experiencing degraded performance. We are continuing to investigate.
  11. Investigating · Aug 26, 15:11 UTC
    We are investigating reports of degraded availability for Actions

More from GitHub

Full history
StartedIncidentDuration
Sep 2310:11 UTCIncident across several servicesOngoing
Sep 2022:13 UTCIncident with Pull Requests1h 9m
Sep 1720:59 UTCElevated rate of errors for OpenAI models provided by Copilot50m
Sep 1607:20 UTCDegradation with Gemini 3.8 Flash10h 28m
Sep 1519:11 UTCDisruption with some GitHub services49m
Sep 1509:47 UTCDisruption with some GitHub services1h 30m

Also caused by capacity and load

All
StartedIncidentDuration
Sep 2213:19 UTCWe are investigating an issue with CH servers in Azure germanywestcentral regionClickHouse5h 6m
Sep 221:44 UTCElevated Linux worker queue timesExpo21m
Aug 2506:53 UTC[Critical] Issue with DownloadsBox44m
Aug 2015:40 UTCWe are investigating an issue where customers may experience timeouts, service degradations, errors, and elevated latencies across multiple products in the us-west1 region.Google Cloud3h 40m
Aug 1812:58 UTCINC20000163Snowflake1h 54m
Aug 1713:19 UTCNo capacity in ARNFly.io2h 22m

From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.

Weekly: the week's major outages, postmortems and breaches, Saturday mornings.