Disruption with some GitHub services
Final update
On July 24th at 16:04 UTC, a loss of connectivity occurred in network paths in one of our three physical data center availability zones (AZs). This resulted in packet loss due to the remaining active paths becoming saturated. Our data centers use a leaf-spine switch fabric in each compute cage, and an aggregation layer interconnecting the spines from each cage within each AZ. The loss of connectivity affected links between one cage’s spine switches and the aggregation layer within that specific AZ. Workloads depending on compute resources in this cage became degraded due to packet loss, and exhibited intermittent errors: - Actions saw 10% of jobs fail during the impact window, and 5% of jobs succeeded but with delayed starts. - 27% of GitHub issues interactions saw slow requests or timeouts. - 4% of GitHub Copilot requests experienced errors, though most automatically retry. - 4% of git push operations saw impacts during the affected window. - Authentication requests saw increased latency during the affected window, but error rates, while elevated, were < 1% in all cases. We were able to mitigate the outage by re-routing affected connections to available fiber paths that wer
Timeline
- Resolved · Jul 24, 17:36 UTC
On July 24th at 16:04 UTC, a loss of connectivity occurred in network paths in one of our three physical data center availability zones (AZs). This resulted in packet loss due to the remaining active paths becoming saturated. Our data centers use a leaf-spine switch fabric in each compute cage, and an aggregation layer interconnecting the spines from each cage within each AZ. The loss of connectivity affected links between one cage’s spine switches and the aggregation layer within that specific AZ. Workloads depending on compute resources in this cage became degraded due to packet loss, and exhibited intermittent errors: - Actions saw 10% of jobs fail during the impact window, and 5% of jobs succeeded but with delayed starts. - 27% of GitHub issues interactions saw slow requests or timeouts. - 4% of GitHub Copilot requests experienced errors, though most automatically retry. - 4% of git push operations saw impacts during the affected window. - Authentication requests saw increased latency during the affected window, but error rates, while elevated, were < 1% in all cases. We were able to mitigate the outage by re-routing affected connections to available fiber paths that were allocated for future capacity upgrades. Sufficient network capacity to eliminate packet loss was restored at 17:07, with most services showing full recovery by 17:16. All paths were restored and services healthy at 17:36. This incident affected 25% of available network interconnect capacity. Olde
- Investigating · Jul 24, 17:24 UTC
We are seeing recovery across all services
- Investigating · Jul 24, 17:16 UTC
The degradation affecting API Requests, Actions, Copilot, Issues, Pages and Pull Requests has been mitigated. We are monitoring to ensure stability.
- Investigating · Jul 24, 16:41 UTC
Actions is experiencing degraded performance. We are continuing to investigate.
- Investigating · Jul 24, 16:40 UTC
We have applied a mitigation and are monitoring for recovery
- Investigating · Jul 24, 16:28 UTC
Actions is experiencing degraded availability. We are continuing to investigate.
- Investigating · Jul 24, 16:27 UTC
Pages is experiencing degraded performance. We are continuing to investigate.
- Investigating · Jul 24, 16:26 UTC
Copilot is experiencing degraded performance. We are continuing to investigate.
- Investigating · Jul 24, 16:22 UTC
We are investigating timeouts to some GitHub services
- Investigating · Jul 24, 16:20 UTC
Pull Requests is experiencing degraded performance. We are continuing to investigate.
- Investigating · Jul 24, 16:19 UTC
Actions is experiencing degraded performance. We are continuing to investigate.
- Investigating · Jul 24, 16:17 UTC
We are investigating reports of degraded performance for API Requests and Issues
More from GitHub
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 2310:11 UTC | Incident across several services | minor | Ongoing |
| Sep 2022:13 UTC | Incident with Pull Requests | minor | 1h 9m |
| Sep 1720:59 UTC | Elevated rate of errors for OpenAI models provided by Copilot | minor | 50m |
| Sep 1607:20 UTC | Degradation with Gemini 3.8 Flash | major | 10h 28m |
| Sep 1519:11 UTC | Disruption with some GitHub services | minor | 49m |
| Sep 1509:47 UTC | Disruption with some GitHub services | minor | 1h 30m |
Also caused by capacity and load
All| Started | Vendor | Incident | Impact | Duration |
|---|---|---|---|---|
| Sep 2213:19 UTC | We are investigating an issue with CH servers in Azure germanywestcentral region | minor | 5h 6m | |
| Sep 221:44 UTC | Elevated Linux worker queue times | minor | 21m | |
| Aug 2506:53 UTC | [Critical] Issue with Downloads | critical | 44m | |
| Aug 2015:40 UTC | We are investigating an issue where customers may experience timeouts, service degradations, errors, and elevated latencies across multiple products in the us-west1 region. | critical | 3h 40m | |
| Aug 1812:58 UTC | INC20000163 | major | 1h 54m | |
| Aug 1713:19 UTC | No capacity in ARN | none | 2h 22m |
From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.