Skip to content
GitHub · Developer toolsJan 30, 2026, 20:59 UTC

Degraded Experience - Failing to finalize some CCA Jobs

MinorConfig changeUpdated 16h ago
Jan 30, 20:59 UTCJan 30, 21:22 UTC
Duration
23m
Impact
Minor
Root cause
Config change
GitHub, 90 days
68 incidents
Affected
Not listed by the vendor.
Status page

Final update

Between 2026-01-30 19:06 UTC and 2026-01-30 20:04 UTC, Copilot Coding Agent experienced sessions getting stuck, with a mismatch between the UI-reported session status and the underlying Actions and job execution state. Impacted users could observe Actions finish successfully but the session UI continuing to show in-progress state, or sessions remaining in queued state. The issue was caused by a feature flag that resulted in events being published to a new Kafka topic. Publishing failures led to buffer/queue overflows in the shared event publishing client, preventing other critical events from being emitted. We mitigated the incident by disabling the feature flag and redeploying production pods, which resumed normal event delivery. We are working to improve safeguards and detection around event publishing failures to reduce time to mitigation for similar issues in the future.

Timeline

  1. Resolved · Jan 30, 21:22 UTC
    Between 2026-01-30 19:06 UTC and 2026-01-30 20:04 UTC, Copilot Coding Agent experienced sessions getting stuck, with a mismatch between the UI-reported session status and the underlying Actions and job execution state. Impacted users could observe Actions finish successfully but the session UI continuing to show in-progress state, or sessions remaining in queued state. The issue was caused by a feature flag that resulted in events being published to a new Kafka topic. Publishing failures led to buffer/queue overflows in the shared event publishing client, preventing other critical events from being emitted. We mitigated the incident by disabling the feature flag and redeploying production pods, which resumed normal event delivery. We are working to improve safeguards and detection around event publishing failures to reduce time to mitigation for similar issues in the future.

More from GitHub

Full history
StartedIncidentDuration
Sep 2310:11 UTCIncident across several servicesOngoing
Sep 2022:13 UTCIncident with Pull Requests1h 9m
Sep 1720:59 UTCElevated rate of errors for OpenAI models provided by Copilot50m
Sep 1607:20 UTCDegradation with Gemini 3.8 Flash10h 28m
Sep 1519:11 UTCDisruption with some GitHub services49m
Sep 1509:47 UTCDisruption with some GitHub services1h 30m

Also caused by configuration change

All
StartedIncidentDuration
Sep 2218:20 UTCPhone Number APIs and Console Were Returning Incorrect 404 ResponsesTwilio0m
Sep 1622:30 UTCHyperdrive Elevated Origin Connection Failure RatesCloudflare0m
Sep 1211:04 UTCSome customers experiencing blurry image previews and download issuesSlack8h 56m
Sep 1107:18 UTCINC20000213Snowflake6h 1m
Sep 407:34 UTCINC20000199Snowflake2h 36m
Sep 109:58 UTCINC20000190Snowflake3h 52m

From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.

Weekly: the week's major outages, postmortems and breaches, Saturday mornings.