Fly.io Upstash Redis service distruption (FRA region)
Final update
On May 12th and 13th at various times, a subset of Upstash Redis instances on [Fly.io](http://Fly.io) experienced intermittent hangs and elevated error rates. The Redis process would stall inside a logging syscall , alive but not making progress , which made the issue hard to spot from our usual telemetry. After investigating with Fly's team, we identified the root cause as a bad interaction between a recent guest kernel update on Fly's newer machines and an upstream Cloud Hypervisor bug \([cloud-hypervisor#7672](https://github.com/cloud-hypervisor/cloud-hypervisor/issues/7672)\) affecting log writes from inside the VM. We mitigated by disabling the affected logging paths, and Fly has since rolled out a hypervisor-side patch, fully resolving the issue. No data was lost. Sorry for the disruption.
Timeline
- Postmortem · May 15, 12:49 UTC
On May 12th and 13th at various times, a subset of Upstash Redis instances on [Fly.io](http://Fly.io) experienced intermittent hangs and elevated error rates. The Redis process would stall inside a logging syscall , alive but not making progress , which made the issue hard to spot from our usual telemetry. After investigating with Fly's team, we identified the root cause as a bad interaction between a recent guest kernel update on Fly's newer machines and an upstream Cloud Hypervisor bug \([cloud-hypervisor#7672](https://github.com/cloud-hypervisor/cloud-hypervisor/issues/7672)\) affecting log writes from inside the VM. We mitigated by disabling the affected logging paths, and Fly has since rolled out a hypervisor-side patch, fully resolving the issue. No data was lost. Sorry for the disruption.
- Resolved · May 11, 18:22 UTC
The incident has been resolved. We are working with Fly team on RCA.
- Investigating · May 11, 17:19 UTC
We are working with Fly team to investigate the root cause.
- Investigating · May 11, 15:07 UTC
We are continuing to investigate the issue.
- Investigating · May 11, 15:05 UTC
Some databases may experience increased latency or timeouts in Fly.io’s FRA region.
More from Upstash
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 913:00 UTC | Intermittent DNS Resolution Errors for Upstash Vector in US East (us-east-1) | minor | 0m |
| Sep 308:29 UTC | Upstash Console login issue | none | 0m |
| Aug 2822:56 UTC | Message Persistence Issue , QStash us-east-1 | major | 0m |
| Jul 2208:50 UTC | Fly.io infrastructure disruption affecting some Upstash Redis databases on Fly.io DFW Region | none | 1h 15m |
| Jul 1608:35 UTC | QStash us-east-1 - URL Publish Errors | none | 0m |
| Jun 2515:23 UTC | Upstash Redis Partial Service Disruption | major | 1h 12m |
Also caused by software bug
All| Started | Vendor | Incident | Impact | Duration |
|---|---|---|---|---|
| Sep 1518:57 UTC | ClickPipes failing on Kinesis in AWS us-east-1 | critical | 29h 46m | |
| Sep 318:20 UTC | Retroactive Incident: Twilio Personalized Support Phone Line Affected | none | 0m | |
| Aug 819:48 UTC | Alerting expressions pipeline failing when recovery settings | minor | 0m | |
| Aug 615:22 UTC | Incident with Actions | critical | 10h 42m | |
| Jul 2314:14 UTC | [Medium] Issues with Box Hubs | major | 16m | |
| Jun 1719:00 UTC | Incident With Webhooks | none | 0m |
From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.