Upstream Azure Outage Causing Service Degradation
May 29, 06:32 UTCMay 31, 21:33 UTC
Duration
2d 15h
Impact
Major
Root cause
Power
Elastic, 90 days
27 incidents
Affected
Azure Washington (azure-westus2) - Elasticsearch connectivity: Azure azure-westus2Azure Washington (azure-westus2) - Kibana connectivity: Azure azure-westus2Azure Washington (azure-westus2) - Azure Infrastructure health: azure-westus2Azure Washington (azure-westus2) - APM connectivity: Azure azure-westus2Azure Washington (azure-westus2) - Deployment orchestration (Create/Edit/Restart/Delete): Azure azure-westus2Azure - Washington (azure-westus2) - Azure Infrastructure Health
Final update
This incident has been resolved.
Timeline
- Resolved · May 31, 21:33 UTC
This incident has been resolved.
- Monitoring · May 30, 06:36 UTC
We continue to monitor the situation which has mostly recovered. A small number of deployments are still not healthy. Our support team is working on resolving these on a case-by-case basis.
- Monitoring · May 29, 18:50 UTC
Remediation is underway and systems are recovering. We continue to monitor the situation and will post an update in the next 2 hours or earlier.
- Identified · May 29, 14:26 UTC
We continue to remediate the impact of an upstream Azure outage in Azure West US 2 (azure-westus2), limited to availability zone westus2-2. Customers with deployments not configured for high availability in this region may continue to experience intermittent connectivity errors, degraded performance, or delays when performing deployment operations. Deployments configured for high availability across multiple availability zones are not expected to be affected. Our teams remain actively engaged with Microsoft Azure and are working to restore affected infrastructure. We will provide further updates as recovery progresses
- Identified · May 29, 11:46 UTC
We continue to respond to an upstream Microsoft Azure outage in Azure West US 2 (azure-westus2), caused by a datacenter power event reported by Azure. Elastic impact remains limited to availability zone westus2-2. Customers with deployments not configured for high availability in this region may continue to experience intermittent connectivity errors, degraded performance, or delays when performing deployment operations (including create, edit, restart, and delete). Deployments configured for high availability across multiple availability zones are not expected to be affected. Our teams are actively remediating affected infrastructure in the impacted zone. Recovery progress is currently constrained by degraded Azure platform APIs and capacity availability in the region.
- Identified · May 29, 10:43 UTC
We are continuing to investigate the impact of an upstream Azure WestUS2 outage which is impacting Elastic services and deployment healthiness, including Elastic Serverless projects hosted in this region. We have confirmed that impact to our service is limited to a single availability zone (westus2-2). Customers whose deployments are not configured for high availability may experience intermittent errors, delays, or inability to access certain features. Deployments configured for high availability across multiple zones are not expected to be affected. Our team is working to mitigate the situation and the impact on our environment. We will provide further updates as more information becomes available.
- Investigating · May 29, 06:50 UTC
We are continuing to investigate this issue.
- Investigating · May 29, 06:35 UTC
We are continuing to investigate this issue.
- Investigating · May 29, 06:32 UTC
We are continuing to investigate the impact of an upstream Azure WestUS2 outage which is impacting Elastic services and deployment healthiness, including Elastic Serverless projects hosted in this region. We have confirmed that impact to our service is limited to a single availability zone (westus2-2). Customers whose deployments are not configured for high availability may experience intermittent errors, delays, or inability to access certain features. Deployments configured for high availability across multiple zones are not expected to be affected. Our team is working to mitigate the situation and the impact on our environment. We will provide further updates as more information becomes available.
More from Elastic
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 2314:04 UTC | Degraded Performance: Cloud Provisioning Delays | major | 1h 59m |
| Sep 1715:50 UTC | AutoOps node metrics temporarily unavailable in some regions | major | 41m |
| Sep 1613:17 UTC | Elastic Support Portal unavailable | major | 2h 21m |
| Sep 1223:21 UTC | Delayed Metrics in Cloud Console - GCP us-east4 | major | 8h 24m |
| Sep 1001:30 UTC | Kibana access restored for UI-assigned Organization Owners on Hosted deployments | major | 21h |
| Sep 920:44 UTC | Elastic Agent enrollment/check-in failures on 9.5.3 (and 9.4.6) with Fleet remote Elasticsearch output | major | 7d 13h |
Also caused by power and facilities
All| Started | Vendor | Incident | Impact | Duration |
|---|---|---|---|---|
| Jul 1523:57 UTC | Google Cloud VMware Engine (GCVE), Google Cloud NetApp Volumes, and Bare Metal Solutions (BMS) services are experiencing a service outage in europe-west4-a due to a cooling failure. | major | 12h 28m | |
| May 2906:47 UTC | Azure - West US 2 (Washington): INC0157956 | major | 15h 42m | |
| May 800:25 UTC | Increased Error Rate and Latency | minor | 26h 39m | |
| Mar 113:33 UTC | AWS - Middle East (UAE): INC0152681 | critical | 75d 1h |
From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.