[Critical] Issues with Multiple Box Services
Final update
We recently addressed issues affecting folder and collection browsing, and related APIs. We would like to take the opportunity to further explain these issues and the steps we have taken to keep them from happening in the future. Between 05:22 PM PDT and 07:04 PM PDT on May 6, 2026, some users may have experienced difficulties while working in Box. During this time, users reported failures and elevated errors when loading folders and collections, intermittent service errors, and degraded AI-related functionality. The issue was caused by runtime saturation in a backend API service that had been gradually approaching its capacity limits over several weeks. During peak traffic hours, multiple service instances simultaneously became unable to respond to internal health checks within the configured timeout. As those instances were restarted, traffic shifted to the remaining instances, increasing load on them and creating a cascading failure. We mitigated the issue by stopping the restart cycle, routing traffic to alternate paths, and increasing service capacity. **Analysis** This incident revealed several contributing causes: * Insufficient scaling signals , autoscaling was primarily
Timeline
- Postmortem · May 15, 17:48 UTC
We recently addressed issues affecting folder and collection browsing, and related APIs. We would like to take the opportunity to further explain these issues and the steps we have taken to keep them from happening in the future. Between 05:22 PM PDT and 07:04 PM PDT on May 6, 2026, some users may have experienced difficulties while working in Box. During this time, users reported failures and elevated errors when loading folders and collections, intermittent service errors, and degraded AI-related functionality. The issue was caused by runtime saturation in a backend API service that had been gradually approaching its capacity limits over several weeks. During peak traffic hours, multiple service instances simultaneously became unable to respond to internal health checks within the configured timeout. As those instances were restarted, traffic shifted to the remaining instances, increasing load on them and creating a cascading failure. We mitigated the issue by stopping the restart cycle, routing traffic to alternate paths, and increasing service capacity. **Analysis** This incident revealed several contributing causes: * Insufficient scaling signals , autoscaling was primarily configured around CPU utilization. For this service, CPU utilization did not accurately reflect runtime saturation, so the fleet did not scale early enough in response to the actual bottleneck. * Health check sensitivity under load , health checks were configured with timeouts that did not account
- Resolved · May 7, 03:07 UTC
After further monitoring, this incident is now considered resolved. If you continue to experience any issues, please contact Box Support at https://support.box.com.
- Monitoring · May 7, 02:32 UTC
Our team has taken steps to remediate this issue and is seeing improvement in Box Folder, Collections, Box AI, Hubs, and Classification. We are continuing to monitor for any additional impact.
- Identified · May 7, 02:12 UTC
Our team has identified the underlying cause of this issue and is working to take remediating steps. We will provide additional updates as they become available.
- Identified · May 7, 01:52 UTC
Our team has identified the underlying cause of this issue and is working to take remediating steps. We will provide additional updates as they become available.
- Investigating · May 7, 01:20 UTC
We are investigating an ongoing issue affecting browsing folders, collections, Box AI, Classifications, and Hubs. Customers may experience seeing errors when loading folders, collections, Box AI and Hubs. We will provide more information as soon as it is available.
More from Box
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 402:10 UTC | [Minor] Issue with Admin Reporting | minor | 53m |
| Aug 2808:20 UTC | [Medium] Issues with Box Hubs Shared Links | minor | 1h 40m |
| Aug 2506:53 UTC | [Critical] Issue with Downloads | critical | 44m |
| Aug 2421:21 UTC | [Medium] Issue with Uploads | major | 41m |
| Aug 2417:30 UTC | [Minor] Issue with Logins & Files Page | minor | 0m |
| Aug 2000:31 UTC | [Medium] Issue with Microsoft Office Integration and Box Edit | major | 1h 13m |
Also caused by capacity and load
All| Started | Vendor | Incident | Impact | Duration |
|---|---|---|---|---|
| Sep 2213:19 UTC | We are investigating an issue with CH servers in Azure germanywestcentral region | minor | 5h 6m | |
| Sep 1607:20 UTC | Degradation with Gemini 3.8 Flash | major | 10h 28m | |
| Sep 1509:47 UTC | Disruption with some GitHub services | minor | 1h 30m | |
| Sep 422:02 UTC | Degradation in repos contents API | minor | 21m | |
| Sep 221:44 UTC | Elevated Linux worker queue times | minor | 21m | |
| Sep 115:00 UTC | Delays in commit processing | minor | 1h 1m |
From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.