Skip to content
GitHub · Developer toolsAug 6, 2026, 15:22 UTC

Incident with Actions

CriticalBugUpdated 8h ago
Aug 6, 15:22 UTCAug 7, 02:04 UTC
Duration
10h 42m
Impact
Critical
Root cause
Bug
GitHub, 90 days
68 incidents
Affected
ActionsPages

Lesson: Teams should verify service health when experiencing disruptions with deployment actions.

What happened

On August 6, 2026, between 15:05 UTC and 00:14 UTC on August 7, GitHub Actions experienced degraded availability. During the incident, workflow runs failed or remained queued for an extended period of time. Customers using both GitHub-hosted and self-hosted runners were affected. At peak, 71% of workflow runs experienced infrastructure failures and 75% of the remaining workflow runs were delayed by more than 5 minutes. The incident was triggered by a routine deployment to an internal Actions service responsible for processing events and generating Actions jobs. The deployment exposed an existing capacity and concurrency weakness. As pods were replaced during the deployment, remaining capacity became saturated, causing services to crash and triggering a cascading impact across multiple clusters and downstream services. These services recovered at 17:00 after expanding capacity, throttling incoming webhook-triggered work to allow the system to recover, and increasing processing capacity for the backlog of affected events. As the incident progressed, a backlog of work accumulated across the systems responsible for assigning jobs to runners. Due to a latent bug in one of the service

Timeline

  1. Resolved · Aug 7, 02:04 UTC
    On August 6, 2026, between 15:05 UTC and 00:14 UTC on August 7, GitHub Actions experienced degraded availability. During the incident, workflow runs failed or remained queued for an extended period of time. Customers using both GitHub-hosted and self-hosted runners were affected. At peak, 71% of workflow runs experienced infrastructure failures and 75% of the remaining workflow runs were delayed by more than 5 minutes. The incident was triggered by a routine deployment to an internal Actions service responsible for processing events and generating Actions jobs. The deployment exposed an existing capacity and concurrency weakness. As pods were replaced during the deployment, remaining capacity became saturated, causing services to crash and triggering a cascading impact across multiple clusters and downstream services. These services recovered at 17:00 after expanding capacity, throttling incoming webhook-triggered work to allow the system to recover, and increasing processing capacity for the backlog of affected events. As the incident progressed, a backlog of work accumulated across the systems responsible for assigning jobs to runners. Due to a latent bug in one of the services responsible for job assignment, runners were getting assigned jobs that were no longer valid and then getting stuck retrying those jobs, preventing them from picking up valid work. This second stage of impact was mitigated by deploying changes to prevent runners from repeatedly attempting to acqu
  2. Monitoring · Aug 7, 02:03 UTC
    During the incident, some Actions Runner Controller (ARC) runner pods became stuck in an idle state. Affected users can delete those pods using kubectl or redeploy their Actions Runner Controller application. ARC will automatically create replacement runners. The next releases of Actions Runner and Actions Runner Controller will include an automatic recovery mechanism, preventing the need for these manual steps in the future. Some workflow-triggering events, including push and pull request events, were not processed during the incident and cannot be replayed automatically. Customers may need to repeat the triggering action by pushing a new commit, updating the pull request, or manually re-running the workflow where applicable.
  3. Monitoring · Aug 7, 00:59 UTC
    We’re investigating reports that some Actions Runner Controller runners are taking longer than expected to recover. We’ll provide an update as our investigation progresses.
  4. Monitoring · Aug 7, 00:06 UTC
    The degradation has been mitigated. We are monitoring to ensure stability.
  5. Investigating · Aug 7, 00:05 UTC
    The degradation affecting Actions and Pages has been mitigated. We are monitoring to ensure stability.
  6. Investigating · Aug 7, 00:01 UTC
    System-wide queues have been drained, and new jobs are being processed as expected. The fix for self-hosted runners not picking up jobs has been fully rolled out. Webhook-triggered Actions workflows have been restored to full throughput. GitHub Pages, Copilot code review, and Copilot coding agent are showing recovery. Migrations using GitHub Enterprise Importer remain paused as a precaution. We are monitoring all affected services for sustained recovery and will provide another update shortly.
  7. Investigating · Aug 7, 00:01 UTC
    System-wide queues have been drained, and new jobs are being processed as expected. The fix for self-hosted runners not picking up jobs has been fully rolled out. Webhook-triggered Actions workflows have been restored to full throughput. GitHub Pages, Copilot code review, and Copilot coding agent are showing recovery. Migrations using GitHub Enterprise Importer remain paused as a precaution. We are monitoring all affected services for sustained recovery and will provide another update shortly.
  8. Investigating · Aug 6, 23:13 UTC
    We have deployed fixes that address runners being assigned invalid jobs and are taking additional steps to clear the backlog of affected jobs. Job completion rates for running workflows have improved significantly, with success rates now at 99%. Global queues for hosted runner assignment are nearly burned down and concurrency queues for customers are being processed. Another change was deployed to accelerate processing the backlog of job requests. We are gradually restoring throughput for webhook-triggered Actions workflows and monitoring system stability. We have deployed a fix for self-hosted runners that were not picking up jobs and are enabling it incrementally. GitHub Pages, Copilot code review, and Copilot coding agent may still experience intermittent failures or delays. Migrations using GitHub Enterprise Importer remain paused. We continue to monitor recovery across all affected services and will provide another update as conditions improve.
  9. Investigating · Aug 6, 22:18 UTC
    We continue to make progress on the issue affecting GitHub Actions. We have deployed a fix that addresses runners being assigned jobs that are no longer valid, and are seeing improvement in job completion rates. For workflow runs that are starting, success rates have increased significantly and are now at 97%. Standard and larger runners are now draining queued work. A change is also in progress to mitigate issues with existing self-hosted runners that are not picking up jobs. Webhook triggers remain throttled to support recovery. Many push and pull request events are not yet triggering new workflow runs, and we are working to safely restore full throughput. GitHub Pages, Copilot code review, and Copilot coding agent may still experience failures or delays. Migrations using GitHub Enterprise Importer remain paused. We are continuing to monitor recovery and will provide another update as conditions improve.
  10. Investigating · Aug 6, 21:30 UTC
    We are continuing to work on an issue affecting GitHub Actions. Webhook triggers remain throttled to aid recovery, so many push and pull request events are not triggering new workflow runs. We identified runners being assigned jobs that are no longer valid and are deploying a change to address this issue. Both GitHub-hosted and self-hosted runners are affected. Copilot code review, Copilot coding agent, and GitHub Pages may experience failures or delays. Migrations using GitHub Enterprise Importer have been paused to support mitigation efforts.
  11. Investigating · Aug 6, 20:34 UTC
    We are continuing to work on an issue affecting GitHub Actions. Webhook triggers are currently throttled to help with recovery and and we are processing approximately 15% of webhooks, so many events such as pushes and pull requests are not triggering workflow runs. Of jobs queued, approximately 65% are succeeding, improved from a low of 30 to 40% earlier in this incident. We have narrowed the remaining impact to runners that are stuck retrying jobs that are no longer available. Both GitHub-hosted and self-hosted runners are affected, and we are working to recover them. Copilot code review, Copilot coding agent, and migrations using GitHub Enterprise Importer may also be affected.
  12. Investigating · Aug 6, 19:43 UTC
    We are continuing to work on an issue affecting GitHub Actions. Capacity remains constrained and jobs may still be delayed or fail while it recovers gradually. Customers using self-hosted runners may see errors or rate limiting when runners register.  Copilot code review, Copilot coding agent, and migrations using GitHub Enterprise Importer may also be affected. Webhook deliveries may be delayed. Our engineers remain actively engaged.
  13. Investigating · Aug 6, 18:46 UTC
    We are continuing to work on an issue affecting multiple GitHub services. Workflow runs are still failing, and jobs may remain queued for an extended period before starting or may time out. Jobs using GitHub-hosted runners are particularly affected while capacity is constrained. Customers using self-hosted runners may see errors or rate limiting when runners register. Copilot code review, Copilot coding agent, and migrations using GitHub Enterprise Importer may also be affected. Webhook deliveries may be delayed. Recovery is taking longer than we expected, and engineers remain actively engaged.
  14. Investigating · Aug 6, 18:11 UTC
    We are continuing to work on an issue affecting multiple GitHub services. Workflow runs are still failing or delayed in starting, and some queued jobs may time out. Customers using self-hosted runners may see errors or rate limiting when runners register. Copilot code review, Copilot coding agent, hosted runners, and migrations using GitHub Enterprise Importer may also be affected. Webhook deliveries may be delayed. Engineers have applied further mitigations and are continuing to work towards full recovery.
  15. Investigating · Aug 6, 17:40 UTC
    We are continuing to work on an issue affecting multiple GitHub services. Workflow runs are failing or delayed in starting, and some queued jobs may time out. Copilot code review, Copilot coding agent, hosted runners, and migrations using GitHub Enterprise Importer might also affected. Webhook deliveries may be delayed. Engineers have applied a number of mitigations and are rolling out a further fix across all affected systems now.
  16. Investigating · Aug 6, 17:02 UTC
    We are continuing to work on the issue affecting GitHub Actions. Workflow runs are still failing or delayed in starting, and some queued jobs may time out. Some requests to the Actions API are returning errors. Customers running migrations with GitHub Enterprise Importer may see failures. Our engineers have applied several mitigations and are rolling out a further fix now.

More from GitHub

Full history
StartedIncidentDuration
Sep 2310:11 UTCIncident across several servicesOngoing
Sep 2022:13 UTCIncident with Pull Requests1h 9m
Sep 1720:59 UTCElevated rate of errors for OpenAI models provided by Copilot50m
Sep 1607:20 UTCDegradation with Gemini 3.8 Flash10h 28m
Sep 1519:11 UTCDisruption with some GitHub services49m
Sep 1509:47 UTCDisruption with some GitHub services1h 30m

Also caused by software bug

All
StartedIncidentDuration
Sep 1518:57 UTCClickPipes failing on Kinesis in AWS us-east-1ClickHouse29h 46m
Sep 318:20 UTCRetroactive Incident: Twilio Personalized Support Phone Line AffectedTwilio0m
Aug 819:48 UTCAlerting expressions pipeline failing when recovery settingsGrafana Labs0m
Jul 2314:14 UTC[Medium] Issues with Box HubsBox16m
May 1115:05 UTCFly.io Upstash Redis service distruption (FRA region)Upstash3h 17m
May 809:46 UTCQStash US Region Service DisruptionUpstash22m

From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.

Weekly: the week's major outages, postmortems and breaches, Saturday mornings.