AutoOps node metrics temporarily unavailable in some regions
Sep 17, 15:50 UTCSep 17, 16:31 UTC
Duration
41m
Impact
Major
Root cause
Storage
Elastic, 90 days
27 incidents
Affected
Not listed by the vendor.
Final update
The issue causing missing node-level metrics (CPU, memory, disk, thread pools) for recent time ranges in a subset of AutoOps regions has been fully resolved. Cluster health, shard, and deployment data were unaffected throughout, and no data was lost.
Timeline
- Resolved · Sep 17, 16:31 UTC
The issue causing missing node-level metrics (CPU, memory, disk, thread pools) for recent time ranges in a subset of AutoOps regions has been fully resolved. Cluster health, shard, and deployment data were unaffected throughout, and no data was lost.
- Identified · Sep 17, 16:13 UTC
We've identified the cause of the missing node-level metrics (CPU, memory, disk, thread pools) in affected regions. Cluster health, shard, and deployment data were never affected, and no data loss has occurred. A fix has been validated in one region and is now being rolled out to the remaining affected regions. We'll provide a further update once the rollout is complete.
- Investigating · Sep 17, 15:50 UTC
We are investigating reports of missing node-level metrics (CPU, memory, disk, thread pools) for recent time ranges in a subset of AutoOps regions. This may appear as a "No data" message on the Nodes view. Cluster health, shard, and deployment data are unaffected, and no data loss has occurred. We will provide an update within the next 2 hours or earlier.
More from Elastic
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 2314:04 UTC | Degraded Performance: Cloud Provisioning Delays | major | 1h 59m |
| Sep 1613:17 UTC | Elastic Support Portal unavailable | major | 2h 21m |
| Sep 1223:21 UTC | Delayed Metrics in Cloud Console - GCP us-east4 | major | 8h 24m |
| Sep 1001:30 UTC | Kibana access restored for UI-assigned Organization Owners on Hosted deployments | major | 21h |
| Sep 920:44 UTC | Elastic Agent enrollment/check-in failures on 9.5.3 (and 9.4.6) with Fleet remote Elasticsearch output | major | 7d 13h |
| Sep 200:04 UTC | Elevated Error Rates for Specific Models Impacting Elastic Inference Service in EU Regions | major | 0m |
Also caused by storage
All| Started | Vendor | Incident | Impact | Duration |
|---|---|---|---|---|
| Sep 1015:26 UTC | Unresponsive Projects | major | 27h 40m | |
| Aug 2623:37 UTC | Disruption with GitHub Billing | minor | 20h 7m | |
| Aug 2413:56 UTC | Actions delays in starting runs | minor | 38m | |
| Aug 1316:21 UTC | Disruption with GHEC Team Sync | minor | 2h 6m | |
| Jul 1909:51 UTC | Block Storage Volume NYC1, NYC3, SGP1, SYD1 and BLR1 | minor | 5h 41m | |
| Jun 1119:42 UTC | Incident with Webhooks | minor | 2h 37m |
From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.