Skip to content
GitHub · Developer toolsOct 1, 2025, 07:59 UTC

Degraded Performance for GitHub Actions MacOS Runners

MinorCapacityUpdated 17h ago
Oct 1, 07:59 UTCOct 1, 16:55 UTC
Duration
8h 56m
Impact
Minor
Root cause
Capacity
GitHub, 90 days
68 incidents
Affected
Not listed by the vendor.
Status page

Final update

On October 1, 2025 between 07:00 UTC and 17:20 UTC, Mac hosted runner capacity for Actions was degraded, leading to timed out jobs and long queue times. On average, the error rate was 46% and peaked at 96% of requests to the service. XL and Intel runners recovered by 10:10 UTC, with the other types taking longer to recover. The degraded capacity was triggered by a scheduled event at 07:00 UTC that led to a permission failure on Mac runner hosts, blocking reimage operations. The permission issue was resolved by 9:41 UTC, but the recovery of available runners took longer than expected due to a combination of backoff logic slowing backend operations and some hosts needing state resets. We deployed changes immediately following the incident to address the scheduled event and ensure that similar failures will not block critical operations in the future. We are also working to reduce the end-to-end time for self-healing of offline hosts for quicker full recovery of future capacity or host events.

Timeline

  1. Resolved · Oct 1, 16:55 UTC
    On October 1, 2025 between 07:00 UTC and 17:20 UTC, Mac hosted runner capacity for Actions was degraded, leading to timed out jobs and long queue times. On average, the error rate was 46% and peaked at 96% of requests to the service. XL and Intel runners recovered by 10:10 UTC, with the other types taking longer to recover. The degraded capacity was triggered by a scheduled event at 07:00 UTC that led to a permission failure on Mac runner hosts, blocking reimage operations. The permission issue was resolved by 9:41 UTC, but the recovery of available runners took longer than expected due to a combination of backoff logic slowing backend operations and some hosts needing state resets. We deployed changes immediately following the incident to address the scheduled event and ensure that similar failures will not block critical operations in the future. We are also working to reduce the end-to-end time for self-healing of offline hosts for quicker full recovery of future capacity or host events.

More from GitHub

Full history
StartedIncidentDuration
Sep 2310:11 UTCIncident across several servicesOngoing
Sep 2022:13 UTCIncident with Pull Requests1h 9m
Sep 1720:59 UTCElevated rate of errors for OpenAI models provided by Copilot50m
Sep 1607:20 UTCDegradation with Gemini 3.8 Flash10h 28m
Sep 1519:11 UTCDisruption with some GitHub services49m
Sep 1509:47 UTCDisruption with some GitHub services1h 30m

Also caused by capacity and load

All
StartedIncidentDuration
Sep 2213:19 UTCWe are investigating an issue with CH servers in Azure germanywestcentral regionClickHouse5h 6m
Sep 221:44 UTCElevated Linux worker queue timesExpo21m
Aug 2506:53 UTC[Critical] Issue with DownloadsBox44m
Aug 2015:40 UTCWe are investigating an issue where customers may experience timeouts, service degradations, errors, and elevated latencies across multiple products in the us-west1 region.Google Cloud3h 40m
Aug 1812:58 UTCINC20000163Snowflake1h 54m
Aug 1713:19 UTCNo capacity in ARNFly.io2h 22m

From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.

Weekly: the week's major outages, postmortems and breaches, Saturday mornings.