Asana full outage for some users (2 hours, 25% outage)
Final update
We’ve been working to add a caching layer to our update pipeline, tuning it carefully and rolling it out gradually. On Wednesday, August 26 we enabled it for most use cases, and it initially performed well. On Monday, August 31, a combination of unrelated infrastructure changes and peak traffic pushed the cache past its scaling limits. Once that threshold was crossed, the cache became unusable, and many pods serving read traffic for the Asana application could no longer serve it. We mitigated the incident by reverting the system to use the previous, non-cached code path. The revert was successful, but recovery took longer than we would expect for this class of issue. A fuller analysis is underway. We will follow up with root causes, action items, and improvements, including why recovery took as long as it did. We were fully down for about 25% of our users, for 2 hours, 15 minutes.
Timeline
- Postmortem · Sep 2, 21:29 UTC
We’ve been working to add a caching layer to our update pipeline, tuning it carefully and rolling it out gradually. On Wednesday, August 26 we enabled it for most use cases, and it initially performed well. On Monday, August 31, a combination of unrelated infrastructure changes and peak traffic pushed the cache past its scaling limits. Once that threshold was crossed, the cache became unusable, and many pods serving read traffic for the Asana application could no longer serve it. We mitigated the incident by reverting the system to use the previous, non-cached code path. The revert was successful, but recovery took longer than we would expect for this class of issue. A fuller analysis is underway. We will follow up with root causes, action items, and improvements, including why recovery took as long as it did. We were fully down for about 25% of our users, for 2 hours, 15 minutes.
- Resolved · Aug 31, 17:18 UTC
User-facing symptoms have recovered. We'll continue to monitor, and will prioritize a retrospective to understand and prevent similar incidents in the future.
- Monitoring · Aug 31, 16:47 UTC
We've applied a fix are seeing signs of recovery.
- Investigating · Aug 31, 16:24 UTC
We have made configuration changes and see partial recovery, but we continue to see some elevated errors.
- Investigating · Aug 31, 15:55 UTC
We are continuing to investigate the issue; we have reverted recent changes, and are working to identify the source of the errors.
- Investigating · Aug 31, 15:14 UTC
We are investigating alerts for slow performance and application errors.
More from Asana
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 2114:59 UTC | Automations not running for EU users | minor | 1h 10m |
| Sep 1714:36 UTC | Slow performance for a subset of Asana users | minor | 36m |
| Sep 218:08 UTC | We are fully down for a subset of customers. We're investigating the issue, and are rolling back a deployment. | major | 42m |
| Aug 415:30 UTC | Webhooks and event streams partial data loss | minor | 0m |
| Jul 718:49 UTC | Errors loading the Asana web app for some users in US | major | 35m |
| Apr 2923:19 UTC | Date Time Trigger Automation Failues | none | 0m |
From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.