Delays starting Docker jobs
Final update
## Summary On Tuesday, June 23, 2026, our engineering team identified that a single Docker job cluster was operating below expected performance levels. Our infrastructure is organized into routing groups, each responsible for handling a specific segment of traffic, and this cluster was the sole member of its group. To address the performance issue, they initiated a standard two-step procedure to safely take it out of rotation: first, update the routing configuration to stop directing traffic to the cluster, then take it offline. When the team moved to take the cluster offline, the first step of the procedure was skipped. The cluster was removed from the available pool, but the routing configuration was not updated, meaning traffic continued to be directed to the routing group even though it had no capacity to process work. With no other clusters in the routing group to absorb the traffic, affected jobs had nowhere to go and began queuing rather than starting. This is a routine and safe operation when the two-step procedure is followed correctly, this incident exposed a gap in our tooling that allowed the procedure to be completed out of order. The issue began at approximately 17:25
Timeline
- Postmortem · Jun 27, 00:15 UTC
## Summary On Tuesday, June 23, 2026, our engineering team identified that a single Docker job cluster was operating below expected performance levels. Our infrastructure is organized into routing groups, each responsible for handling a specific segment of traffic, and this cluster was the sole member of its group. To address the performance issue, they initiated a standard two-step procedure to safely take it out of rotation: first, update the routing configuration to stop directing traffic to the cluster, then take it offline. When the team moved to take the cluster offline, the first step of the procedure was skipped. The cluster was removed from the available pool, but the routing configuration was not updated, meaning traffic continued to be directed to the routing group even though it had no capacity to process work. With no other clusters in the routing group to absorb the traffic, affected jobs had nowhere to go and began queuing rather than starting. This is a routine and safe operation when the two-step procedure is followed correctly, this incident exposed a gap in our tooling that allowed the procedure to be completed out of order. The issue began at approximately 17:25 UTC and was fully resolved by 19:46 UTC, for a total impact window of roughly 2 hours and 20 minutes. Only Docker jobs were affected; other job types, including machine and macOS jobs, continued to operate normally. Beyond the missed step, several factors extended the impact window. An automated al
- Resolved · Jun 23, 19:58 UTC
The issue where customers running Docker jobs experienced significant delays, with jobs stuck in a "preparing environment" state, has now been resolved. We thank you for your patience while our team worked on implementing a fix.
- Monitoring · Jun 23, 19:29 UTC
What's happening Our engineers have deployed a fix and we are observing the job backlog clearing. Docker jobs are now starting normally. What's impacted Customers who submitted Docker jobs during the incident window may still see some jobs completing with elevated queue times. Newly submitted jobs should be starting at normal times. What can you expect Docker job startup times are returning to normal. Thank you for your patience while our engineers resolved this issue. We are monitoring the recovery and will provide a further update if the situation changes.
- Identified · Jun 23, 19:04 UTC
What's impacted Customers running Docker jobs. Affected jobs are queuing rather than starting, resulting in elevated wait times. What can you expect Docker jobs may remain stuck in "preparing environment" for an extended period. Our engineers have identified the cause and are actively working to restore normal job startup times. Thank you for your patience while our team deploys a fix. Next update We will provide another update as soon as more information is available
- Identified · Jun 23, 18:42 UTC
We are deploying a change to mitigate this issue.
- Investigating · Jun 23, 18:34 UTC
We are investigating reports of delays starting Docker jobs.
More from CircleCI
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 2318:13 UTC | Delays in processing new signups and plan changes | none | 1h 53m |
| Sep 2116:22 UTC | Delays starting Linux Machine jobs | minor | 47m |
| Sep 1614:55 UTC | Customers may experience delays in UI updates | minor | 4h 35m |
| Sep 1514:42 UTC | Customers may experience delays in UI updates | minor | 2h 20m |
| Sep 1411:59 UTC | Increased wait times for macOS jobs | minor | 1h |
| Sep 1020:25 UTC | Issues loading pipelines for a small number of projects | minor | 1m |
From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.