Skip to content
Upstash · Data and observabilityMay 11, 2026, 15:05 UTC

Fly.io Upstash Redis service distruption (FRA region)

MajorBugUpdated 19h ago
May 11, 15:05 UTCMay 11, 18:22 UTC
Duration
3h 17m
Impact
Major
Root cause
Bug
Upstash, 90 days
5 incidents
Affected
Not listed by the vendor.
Status page

Final update

On May 12th and 13th at various times, a subset of Upstash Redis instances on [Fly.io](http://Fly.io) experienced intermittent hangs and elevated error rates. The Redis process would stall inside a logging syscall , alive but not making progress , which made the issue hard to spot from our usual telemetry. After investigating with Fly's team, we identified the root cause as a bad interaction between a recent guest kernel update on Fly's newer machines and an upstream Cloud Hypervisor bug \([cloud-hypervisor#7672](https://github.com/cloud-hypervisor/cloud-hypervisor/issues/7672)\) affecting log writes from inside the VM. We mitigated by disabling the affected logging paths, and Fly has since rolled out a hypervisor-side patch, fully resolving the issue. No data was lost. Sorry for the disruption.

Timeline

  1. Postmortem · May 15, 12:49 UTC
    On May 12th and 13th at various times, a subset of Upstash Redis instances on [Fly.io](http://Fly.io) experienced intermittent hangs and elevated error rates. The Redis process would stall inside a logging syscall , alive but not making progress , which made the issue hard to spot from our usual telemetry. After investigating with Fly's team, we identified the root cause as a bad interaction between a recent guest kernel update on Fly's newer machines and an upstream Cloud Hypervisor bug \([cloud-hypervisor#7672](https://github.com/cloud-hypervisor/cloud-hypervisor/issues/7672)\) affecting log writes from inside the VM. We mitigated by disabling the affected logging paths, and Fly has since rolled out a hypervisor-side patch, fully resolving the issue. No data was lost. Sorry for the disruption.
  2. Resolved · May 11, 18:22 UTC
    The incident has been resolved. We are working with Fly team on RCA.
  3. Investigating · May 11, 17:19 UTC
    We are working with Fly team to investigate the root cause.
  4. Investigating · May 11, 15:07 UTC
    We are continuing to investigate the issue.
  5. Investigating · May 11, 15:05 UTC
    Some databases may experience increased latency or timeouts in Fly.io’s FRA region.

More from Upstash

Full history

Also caused by software bug

All
StartedIncidentDuration
Sep 1518:57 UTCClickPipes failing on Kinesis in AWS us-east-1ClickHouse29h 46m
Sep 318:20 UTCRetroactive Incident: Twilio Personalized Support Phone Line AffectedTwilio0m
Aug 819:48 UTCAlerting expressions pipeline failing when recovery settingsGrafana Labs0m
Aug 615:22 UTCIncident with ActionsGitHub10h 42m
Jul 2314:14 UTC[Medium] Issues with Box HubsBox16m
Jun 1719:00 UTCIncident With WebhooksGitHub0m

From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.

Weekly: the week's major outages, postmortems and breaches, Saturday mornings.