Alerting expressions pipeline failing when recovery settings
Aug 8, 19:48 UTCAug 8, 19:48 UTC
Duration
0m
Impact
Minor
Root cause
Bug
Grafana Labs, 90 days
75 incidents
Affected
AWS Australia - prod-ap-southeast-2: Rule EvaluationAWS Australia - prod-ap-southeast-2: Alertmanager and Rules Configuration APIAWS Australia - prod-ap-southeast-2: AlertmanagerAWS Brazil - prod-sa-east-1: Rule EvaluationAWS Brazil - prod-sa-east-1: Alertmanager and Rules Configuration APIAWS Brazil - prod-sa-east-1: AlertmanagerAWS Canada - prod-ca-east-0: Rule EvaluationAWS Canada - prod-ca-east-0: Alertmanager and Rules Configuration APIAWS Canada - prod-ca-east-0: AlertmanagerAWS Germany - prod-eu-west-2: Rule EvaluationAWS Germany - prod-eu-west-2: Alertmanager and Rules Configuration APIAWS Germany - prod-eu-west-2: AlertmanagerAWS UAE - prod-me-central-1: Rule EvaluationAWS UAE - prod-me-central-1: Alertmanager and Rules Configuration API
Final update
Due to a software bug the evaluation of some alert rules (primarily ones that have recovery threshold setting) were failing to be evaluated starting 14:00 UTC to 19:30 UTC today.
Timeline
- Resolved · Aug 8, 19:48 UTC
Due to a software bug the evaluation of some alert rules (primarily ones that have recovery threshold setting) were failing to be evaluated starting 14:00 UTC to 19:30 UTC today.
More from Grafana Labs
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 2221:27 UTC | Kubernetes Observability Billing & Usage Incorrect | minor | 16h 39m |
| Sep 2214:57 UTC | IRM Access Issues for a Small Group of Users | major | 20h 10m |
| Sep 2119:16 UTC | Mimir write request errors | none | 0m |
| Sep 1712:34 UTC | Elevated Latency Managing Cloud Provider Integrations in GCP US Central | minor | 0m |
| Sep 1510:00 UTC | Increased Execution Time for Browser Checks in Synthetic Monitoring | minor | 1h 17m |
| Sep 1507:49 UTC | Intermittent Metric Write Errors in GCP US Central (prod-us-central-0) | minor | 0m |
Also caused by software bug
All| Started | Vendor | Incident | Impact | Duration |
|---|---|---|---|---|
| Sep 1518:57 UTC | ClickPipes failing on Kinesis in AWS us-east-1 | critical | 29h 46m | |
| Sep 318:20 UTC | Retroactive Incident: Twilio Personalized Support Phone Line Affected | none | 0m | |
| Aug 615:22 UTC | Incident with Actions | critical | 10h 42m | |
| Jul 2314:14 UTC | [Medium] Issues with Box Hubs | major | 16m | |
| Jun 1719:00 UTC | Incident With Webhooks | none | 0m | |
| May 2801:13 UTC | Webhook APIs and UI Degraded | minor | 19m |
From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.