Incident with several GitHub Services
Final update
On September 13, 2026, between 08:43 and 10:44 UTC, GitHub experienced degraded availability across approximately 28 services, including Issues, Pull Requests, Actions, Codespaces, Pages, Notifications, Code Scanning, Git LFS, and new account signup. At peak, 8.8% of requests to create GitHub App installation access tokens failed. Token issuance for Actions workflows was also affected, impacting approximately 4% of workflows during the incident time frame. Creating issues through the web interface failed for about 96% of attempts, and signup failures were above 90%. The cause was an internal data-cleanup job that began writing to a shared database cluster at 07:33 UTC. That cluster stores permission data read on nearly every authenticated request. The safeguard that was pacing the background job watched only one health signal , how far the database replicas were lagging , and that signal stayed low the whole time. It did not account for the load building on the primary itself, so the job kept writing while the primary quietly ran toward its limit. When the primary ran out of available connections, requests that needed it could not complete. First, there was no quick timeout on
Timeline
- Resolved · Sep 13, 10:44 UTC
On September 13, 2026, between 08:43 and 10:44 UTC, GitHub experienced degraded availability across approximately 28 services, including Issues, Pull Requests, Actions, Codespaces, Pages, Notifications, Code Scanning, Git LFS, and new account signup. At peak, 8.8% of requests to create GitHub App installation access tokens failed. Token issuance for Actions workflows was also affected, impacting approximately 4% of workflows during the incident time frame. Creating issues through the web interface failed for about 96% of attempts, and signup failures were above 90%. The cause was an internal data-cleanup job that began writing to a shared database cluster at 07:33 UTC. That cluster stores permission data read on nearly every authenticated request. The safeguard that was pacing the background job watched only one health signal , how far the database replicas were lagging , and that signal stayed low the whole time. It did not account for the load building on the primary itself, so the job kept writing while the primary quietly ran toward its limit. When the primary ran out of available connections, requests that needed it could not complete. First, there was no quick timeout on these database calls, so request handlers waited on the stalled database instead of failing fast, and the shared request-handling capacity degraded into site-wide errors. Second, a retry loop around token creation kept re-sending the writes that were already failing, which held the database saturate
- Investigating · Sep 13, 10:28 UTC
Pull Requests is experiencing degraded performance. We are continuing to investigate.
- Investigating · Sep 13, 10:26 UTC
We have reduced load on this cluster with internal load-shedding and are seeing signs of recovery but continue to monitor
- Investigating · Sep 13, 09:36 UTC
We're seeing increased database replication delays on collab which is causing increased error rates in authorization endpoints and follow-on increased error rates across the system - we are investigating
- Investigating · Sep 13, 09:25 UTC
Actions is experiencing degraded performance. We are continuing to investigate.
- Investigating · Sep 13, 09:16 UTC
We are investigating reports of degraded availability for API Requests, Issues, Pages and Pull Requests
More from GitHub
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 2310:11 UTC | Incident across several services | minor | Ongoing |
| Sep 2022:13 UTC | Incident with Pull Requests | minor | 1h 9m |
| Sep 1720:59 UTC | Elevated rate of errors for OpenAI models provided by Copilot | minor | 50m |
| Sep 1607:20 UTC | Degradation with Gemini 3.8 Flash | major | 10h 28m |
| Sep 1519:11 UTC | Disruption with some GitHub services | minor | 49m |
| Sep 1509:47 UTC | Disruption with some GitHub services | minor | 1h 30m |
Also caused by database
All| Started | Vendor | Incident | Impact | Duration |
|---|---|---|---|---|
| Sep 921:53 UTC | INC20000211 | critical | 1h 58m | |
| Aug 2616:17 UTC | Issues performance is degraded in the US | minor | 4h 19m | |
| Aug 818:57 UTC | Account Registration, Droplets, and Related Services | minor | 6h 3m | |
| Jul 2817:41 UTC | Event ingestion and retrieval delays for some North America customers | none | 2h | |
| Jul 300:11 UTC | Partial outage in ORD | major | 5h 48m | |
| Jun 1920:13 UTC | Snyk Code (SAST) Scan service degraded | minor | 2h |
From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.