Performance issues and outages with Cloud products
Final update
### **SUMMARY** We understand the importance of providing reliable and consistent service to our valued customers. On July 6, 2023, from 03:52 to 15:11 UTC, we experienced an issue with an upgraded version of a third-party tool that functions as our internal artifact management system. Despite our monitoring system identifying the incident within two minutes, this issue led to the degradation of the scaling capabilities of our internal hosting platform, resulting in service degradation or outages for customers of Atlassian cloud. In response to this situation, we are taking immediate measures to enhance the stability of our system and prevent similar issues from re-occurring. ### **IMPACT** This incident affected multiple regions and products due to the diminished scaling capabilities of our internal hosting platform. In most products and offerings, customers faced reduced functionality, slower response times, and limited access to specific features. ### **ROOT CAUSE** The root cause of the incident was the introduction of new functionality in a third-party tool that functions as our internal artifact management system. It led to an unexpected increase in the load on the primary da
Timeline
- Postmortem · Jul 14, 04:56 UTC
### **SUMMARY** We understand the importance of providing reliable and consistent service to our valued customers. On July 6, 2023, from 03:52 to 15:11 UTC, we experienced an issue with an upgraded version of a third-party tool that functions as our internal artifact management system. Despite our monitoring system identifying the incident within two minutes, this issue led to the degradation of the scaling capabilities of our internal hosting platform, resulting in service degradation or outages for customers of Atlassian cloud. In response to this situation, we are taking immediate measures to enhance the stability of our system and prevent similar issues from re-occurring. ### **IMPACT** This incident affected multiple regions and products due to the diminished scaling capabilities of our internal hosting platform. In most products and offerings, customers faced reduced functionality, slower response times, and limited access to specific features. ### **ROOT CAUSE** The root cause of the incident was the introduction of new functionality in a third-party tool that functions as our internal artifact management system. It led to an unexpected increase in the load on the primary database of the artifact system. Upon identifying and localizing the problem, we promptly adjusted the system configuration to regain stability. ### **REMEDIAL ACTIONS PLAN & NEXT STEPS** Over the next months, we will enact a temporary freeze on non-critical upgrades of the artifact management system,
- Resolved · Jul 6, 15:37 UTC
We experienced performance issues and outages for several Atlassian Cloud Products. The issue has been resolved and the service is operating normally.
- Monitoring · Jul 6, 13:17 UTC
We have identified the root cause of an issue with an internal infrastructure component that has been impacting multiple Cloud products, including Jira Software, Jira Service Management and Confluence, and customers. This issue had lead to a performance impact and, in some cases, outages. We have implemented a fix to resolve the issue and recovery is in progress.
- Identified · Jul 6, 11:18 UTC
We are investigating an issue with an internal infrastructure component that is impacting multiple Cloud products, including Jira Software, Bitbucket, Jira Service Management and Confluence, and customers. These issues include performance impact and, in some cases, outages. Users may experience slow loading and uploading of attachments, login issues or inability for new customers to sign up. We have identified the root cause and are actively working on the service recovery.
More from Atlassian
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Mar 407:31 UTC | Issues with 403 user authentication errors across Atlassian products | none | 0m |
| Sep 1010:26 UTC | Slow loading times and extended response times for Cloud products in the EU-West region. | none | 2h 18m |
| Sep 908:56 UTC | High RDS CPU on multiple environments | none | 9h 16m |
| Sep 209:23 UTC | Slowness in Jira. | none | 8d 3h |
| Aug 2011:46 UTC | Data residency migrations halted. | minor | 2d 23h |
| Jul 320:51 UTC | Some products are hard down | none | 7h 3m |
Also caused by third-party dependency
All| Started | Vendor | Incident | Impact | Duration |
|---|---|---|---|---|
| Sep 1609:48 UTC | Support Portal Service Disruption | minor | 11h 21m | |
| Sep 1608:31 UTC | Support Helpdesk Availability Issues | minor | 6h 27m | |
| Sep 417:03 UTC | Degradation of Hosted Grafana in US Central Region | major | 3h 39m | |
| Sep 416:05 UTC | Elevated Linux worker queue times | major | 15h 4m | |
| Sep 115:25 UTC | Investigating issues in US Central (prod-us-central-0, prod-us-central-5) | minor | 4h 54m | |
| Aug 3106:26 UTC | Packet loss in ORD | minor | 1h 44m |
From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.