We are fully down for a subset of customers. We're investigating the issue, and are rolling back a deployment.
Final update
Incident: An internal system responsible for automatically scaling backend server capacity in one of our compute clusters stopped replacing capacity that had been cycled out during routine maintenance. This caused a gradual reduction in available capacity over several hours. A subsequent deployment was activated with insufficient capacity, causing all new requests for that compute cluster to fail. To mitigate the impact, we reverted to a previous release revision, which still had sufficient capacity. Impact: For approximately 30 minutes, customers whose traffic was handled by this compute cluster saw full downtime; other customers were unaffected. No customer data was lost. Moving forward: We have added additional monitoring to detect this type of capacity-scaling failure much earlier, and have added safeguards to prevent deployments from shifting traffic before sufficient healthy capacity is confirmed. _Our metric considers a weighted average of uptime experienced by users at each data center. The number of minutes of downtime shown reflects this weighted average._
Timeline
- Postmortem · Sep 4, 19:52 UTC
Incident: An internal system responsible for automatically scaling backend server capacity in one of our compute clusters stopped replacing capacity that had been cycled out during routine maintenance. This caused a gradual reduction in available capacity over several hours. A subsequent deployment was activated with insufficient capacity, causing all new requests for that compute cluster to fail. To mitigate the impact, we reverted to a previous release revision, which still had sufficient capacity. Impact: For approximately 30 minutes, customers whose traffic was handled by this compute cluster saw full downtime; other customers were unaffected. No customer data was lost. Moving forward: We have added additional monitoring to detect this type of capacity-scaling failure much earlier, and have added safeguards to prevent deployments from shifting traffic before sufficient healthy capacity is confirmed. _Our metric considers a weighted average of uptime experienced by users at each data center. The number of minutes of downtime shown reflects this weighted average._
- Resolved · Sep 2, 18:49 UTC
This incident has been resolved.
- Monitoring · Sep 2, 18:14 UTC
A fix has been implemented and we are monitoring the results.
- Investigating · Sep 2, 18:14 UTC
The rollback seems to have completely resolved errors. We're monitoring closely, but users should be able to access Asana again.
- Investigating · Sep 2, 18:13 UTC
We are continuing to investigate this issue.
- Investigating · Sep 2, 18:08 UTC
We are currently investigating this issue.
More from Asana
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 2114:59 UTC | Automations not running for EU users | minor | 1h 10m |
| Sep 1714:36 UTC | Slow performance for a subset of Asana users | minor | 36m |
| Aug 3115:14 UTC | Asana full outage for some users (2 hours, 25% outage) | major | 2h 4m |
| Aug 415:30 UTC | Webhooks and event streams partial data loss | minor | 0m |
| Jul 718:49 UTC | Errors loading the Asana web app for some users in US | major | 35m |
| Apr 2923:19 UTC | Date Time Trigger Automation Failues | none | 0m |
From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.