Grafana Cloud Metrics - Intermittent Write Latency in prod-us-central, prod-us-central-5, and prod-eu-west-0
Final update
This incident is now resolved. During the incident the Cloud Metrics platform experienced intermittent latency spikes communicating with a backend cloud service in the prod-us-central-0 and prod-us-central-5 regions. During the incident the internal CSP-facing issue was escalated to a P1. After determining the scope of the latency spikes was limited to only one availability zone, the team mitigated the situation by migrating all write traffic from to the single nearly unaffected availability zone. As the CSP service team attempted to remedy the situation, the situation became worse and began affecting the previously unaffected zone. Given this, another mitigation path was needed. Changing the connection strategy employed by Cloud Metrics to a different method was deployed to all environments, stabilizing the write path once again as we found the different connection method was more reliable and not affected by these increases in latency. We have migrated all tenants back to multi-zone write paths and are happy with and confident in the current method of connectivity to the backend cloud service, which is the one we migrated to during the course of the incident. We have no immediate
Timeline
- Resolved · Mar 17, 18:22 UTC
This incident is now resolved. During the incident the Cloud Metrics platform experienced intermittent latency spikes communicating with a backend cloud service in the prod-us-central-0 and prod-us-central-5 regions. During the incident the internal CSP-facing issue was escalated to a P1. After determining the scope of the latency spikes was limited to only one availability zone, the team mitigated the situation by migrating all write traffic from to the single nearly unaffected availability zone. As the CSP service team attempted to remedy the situation, the situation became worse and began affecting the previously unaffected zone. Given this, another mitigation path was needed. Changing the connection strategy employed by Cloud Metrics to a different method was deployed to all environments, stabilizing the write path once again as we found the different connection method was more reliable and not affected by these increases in latency. We have migrated all tenants back to multi-zone write paths and are happy with and confident in the current method of connectivity to the backend cloud service, which is the one we migrated to during the course of the incident. We have no immediate plans to use the previous problematic connectivity method for the foreseeable future.
More from Grafana Labs
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 2221:27 UTC | Kubernetes Observability Billing & Usage Incorrect | minor | 16h 39m |
| Sep 2214:57 UTC | IRM Access Issues for a Small Group of Users | major | 20h 10m |
| Sep 2119:16 UTC | Mimir write request errors | none | 0m |
| Sep 1712:34 UTC | Elevated Latency Managing Cloud Provider Integrations in GCP US Central | minor | 0m |
| Sep 1510:00 UTC | Increased Execution Time for Browser Checks in Synthetic Monitoring | minor | 1h 17m |
| Sep 1507:49 UTC | Intermittent Metric Write Errors in GCP US Central (prod-us-central-0) | minor | 0m |
From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.