Skip to content
GitHub · Developer toolsAug 17, 2026, 13:40 UTC

Incident with GitHub.com

CriticalConfig changeUpdated 7h ago
Aug 17, 13:40 UTCAug 17, 21:15 UTC
Duration
7h 36m
Impact
Critical
Root cause
Config change
GitHub, 90 days
68 incidents
Affected
Git OperationsWebhooksAPI RequestsIssuesPull RequestsActionsPagesCopilot

Lesson: Engineering teams should subscribe to platform status updates to remain informed during service incidents.

What happened

On August 17, 2026, from 13:28–21:15 UTC (7h 47m), GitHub.com experienced elevated errors and latency across Issues, Pull Requests, APIs, Actions, and Copilot. At peak, web/API error rates were approximately 20%, while archive and raw-content downloads reached approximately 50%. SAML/OIDC authentication, SCIM, and Team Sync were also affected, as well as Actions workflows in GHEC with Data Residency that depend on public workflow step definitions hosted on GitHub.com. Most services recovered by 16:36 UTC as our Central US datacenter recovered; Actions was degraded until approximately 18:03 UTC; and Copilot Token Service fully recovered by 21:02. Some of the failing traffic was moved from Central US to Northern Virginia where it was served successfully until the network failure in Central US was debugged and resolved. Delayed replies to a single internal endpoint triggered a latent retry bug in VS Code that amplified traffic by approximately 10x and caused delayed recovery for the Copilot Token Service. The immediate cause of the failure was network saturation on load balancers in Central US due to a new peak in traffic. Originally this was caused by an Istio sidecar pod reaching

Timeline

  1. Resolved · Aug 17, 21:15 UTC
    On August 17, 2026, from 13:28–21:15 UTC (7h 47m), GitHub.com experienced elevated errors and latency across Issues, Pull Requests, APIs, Actions, and Copilot. At peak, web/API error rates were approximately 20%, while archive and raw-content downloads reached approximately 50%. SAML/OIDC authentication, SCIM, and Team Sync were also affected, as well as Actions workflows in GHEC with Data Residency that depend on public workflow step definitions hosted on GitHub.com. Most services recovered by 16:36 UTC as our Central US datacenter recovered; Actions was degraded until approximately 18:03 UTC; and Copilot Token Service fully recovered by 21:02. Some of the failing traffic was moved from Central US to Northern Virginia where it was served successfully until the network failure in Central US was debugged and resolved. Delayed replies to a single internal endpoint triggered a latent retry bug in VS Code that amplified traffic by approximately 10x and caused delayed recovery for the Copilot Token Service. The immediate cause of the failure was network saturation on load balancers in Central US due to a new peak in traffic. Originally this was caused by an Istio sidecar pod reaching its concurrency limits and failing to auto scale correctly because of a misconfigured policy that watched host service but not sidecar limits. One failure cascaded to more and eventually four HAProxy nodes exhausted their flow limits, degrading the gateway auth path and causing widespread authentica
  2. Investigating · Aug 17, 20:45 UTC
    We are continuing to apply mitigations to address sporadic Copilot authentication failures in some applications. We expect full recovery within the next 30 minutes. Copilot usage via the GitHub CLI and GitHub App are unaffected.
  3. Investigating · Aug 17, 20:22 UTC
    Issues is operating normally.
  4. Investigating · Aug 17, 20:08 UTC
    We are continuing to investigate sporadic failures affecting Copilot authentication in some applications. Copilot usage via the GitHub CLI and GitHub App are unaffected.
  5. Investigating · Aug 17, 19:13 UTC
    We are continuing to investigate sporadic authentication failures. We have partially disabled authentication token retries and have seen improvement, and we are monitoring impact before fully applying this mitigation.
  6. Investigating · Aug 17, 19:01 UTC
    API Requests is operating normally.
  7. Investigating · Aug 17, 18:48 UTC
    API Requests is experiencing degraded availability. We are continuing to investigate.
  8. Investigating · Aug 17, 18:23 UTC
    The degradation affecting Git Operations has been mitigated. We are monitoring to ensure stability.
  9. Investigating · Aug 17, 18:11 UTC
    We identified the problematic component and have taken corrective actions, but we are seeing residual impact in the form of sporadic authentication failures. We are continuing to apply additional mitigations and investigate the remaining impact.
  10. Investigating · Aug 17, 17:36 UTC
    Issues is experiencing degraded performance. We are continuing to investigate.
  11. Investigating · Aug 17, 17:34 UTC
    We identified the problematic component and have taken corrective actions, but we are seeing residual impact across numerous services. We are continuing to apply additional mitigations and investigate the remaining impact.
  12. Investigating · Aug 17, 17:30 UTC
    Git Operations is experiencing degraded performance. We are continuing to investigate.
  13. Investigating · Aug 17, 16:59 UTC
    The degradation affecting API Requests, Actions, Git Operations, Issues, Pages, Pull Requests and Webhooks has been mitigated. We are monitoring to ensure stability.
  14. Investigating · Aug 17, 16:36 UTC
    We identified the problematic component and have taken corrective actions. There are strong signs of recovery but we are still working to completely restore service, with error rates still remaining slightly elevated. We will post further updates as recovery continues.
  15. Investigating · Aug 17, 16:16 UTC
    We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are still working to identify the root cause and will continue to post updates as we learn more and perform mitigation.
  16. Investigating · Aug 17, 15:42 UTC
    We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are currently performing mitigations and will post updates as we progress.

More from GitHub

Full history
StartedIncidentDuration
Sep 2310:11 UTCIncident across several servicesOngoing
Sep 2022:13 UTCIncident with Pull Requests1h 9m
Sep 1720:59 UTCElevated rate of errors for OpenAI models provided by Copilot50m
Sep 1607:20 UTCDegradation with Gemini 3.8 Flash10h 28m
Sep 1519:11 UTCDisruption with some GitHub services49m
Sep 1509:47 UTCDisruption with some GitHub services1h 30m

Also caused by configuration change

All
StartedIncidentDuration
Sep 2218:20 UTCPhone Number APIs and Console Were Returning Incorrect 404 ResponsesTwilio0m
Sep 1622:30 UTCHyperdrive Elevated Origin Connection Failure RatesCloudflare0m
Sep 1211:04 UTCSome customers experiencing blurry image previews and download issuesSlack8h 56m
Sep 1107:18 UTCINC20000213Snowflake6h 1m
Sep 407:34 UTCINC20000199Snowflake2h 36m
Sep 109:58 UTCINC20000190Snowflake3h 52m

From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.

Weekly: the week's major outages, postmortems and breaches, Saturday mornings.