Skip to content
GitHub · Developer toolsMay 20, 2026, 16:58 UTC

Incident with Actions

MinorConfig changeUpdated 18h ago
May 20, 16:58 UTCMay 20, 20:14 UTC
Duration
3h 16m
Impact
Minor
Root cause
Config change
GitHub, 90 days
68 incidents
Affected
Not listed by the vendor.
Status page

Final update

On May 20, 2026, between 16:00 UTC and 17:45 UTC, GitHub Actions customers experienced run start delays exceeding 5 minutes. Approximately 4.5% of all runs were delayed during the impact window, with scale set jobs disproportionately affected. 30% of scale set jobs were delayed and 4% failed to start entirely. The incident was caused by a misconfigured health check on an internal service that assigns jobs to runners. A brief latency spike in an upstream dependency triggered health check failures across several pods, removing them from service and concentrating load on the remaining capacity. The added load drove memory pressure that escalated into a cascading failure in one regional cluster, leaving it unable to self-recover. Responders mitigated the incident by scaling capacity in the healthy regional clusters and draining traffic away from the impaired one, after which run start latency recovered. To prevent recurrence, we are strengthening our health check configuration to avoid cascading failure scenarios and evaluating automated mitigations to rebalance traffic when a region is degraded.

Timeline

  1. Resolved · May 20, 20:14 UTC
    On May 20, 2026, between 16:00 UTC and 17:45 UTC, GitHub Actions customers experienced run start delays exceeding 5 minutes. Approximately 4.5% of all runs were delayed during the impact window, with scale set jobs disproportionately affected. 30% of scale set jobs were delayed and 4% failed to start entirely. The incident was caused by a misconfigured health check on an internal service that assigns jobs to runners. A brief latency spike in an upstream dependency triggered health check failures across several pods, removing them from service and concentrating load on the remaining capacity. The added load drove memory pressure that escalated into a cascading failure in one regional cluster, leaving it unable to self-recover. Responders mitigated the incident by scaling capacity in the healthy regional clusters and draining traffic away from the impaired one, after which run start latency recovered. To prevent recurrence, we are strengthening our health check configuration to avoid cascading failure scenarios and evaluating automated mitigations to rebalance traffic when a region is degraded.

More from GitHub

Full history
StartedIncidentDuration
Sep 2310:11 UTCIncident across several servicesOngoing
Sep 2022:13 UTCIncident with Pull Requests1h 9m
Sep 1720:59 UTCElevated rate of errors for OpenAI models provided by Copilot50m
Sep 1607:20 UTCDegradation with Gemini 3.8 Flash10h 28m
Sep 1519:11 UTCDisruption with some GitHub services49m
Sep 1509:47 UTCDisruption with some GitHub services1h 30m

Also caused by configuration change

All
StartedIncidentDuration
Sep 2218:20 UTCPhone Number APIs and Console Were Returning Incorrect 404 ResponsesTwilio0m
Sep 1622:30 UTCHyperdrive Elevated Origin Connection Failure RatesCloudflare0m
Sep 1211:04 UTCSome customers experiencing blurry image previews and download issuesSlack8h 56m
Sep 1107:18 UTCINC20000213Snowflake6h 1m
Sep 407:34 UTCINC20000199Snowflake2h 36m
Sep 109:58 UTCINC20000190Snowflake3h 52m

From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.

Weekly: the week's major outages, postmortems and breaches, Saturday mornings.