QStash: Degraded performance in request processing and event logs
Final update
### Product: QStash ### Incident Summary Due to high load, the volume of QStash event logs reached to a point which caused latency in the underlying data store operations. Event log creation was slowed down and lead to performance degradation in QStash request processing. In order to resolve the performance degradation in QStash requests, event logging module was turned off temporarily. After deploying a hot fix and configuration changes, we eventually turned on event logging and system went back to stable state again. ### Root Cause At 07:15 UTC we received alerts on the performance degradation and started the investigation. We discovered long running queries for synching event logs from main QStash servers to QStash event server. In order to resolve performance degradation, we turned off event logging functionality as an immediate action. This action turned the performance back to normal levels for QStash requests but left event log processing disabled. We deployed a hotfix during the day to remove some redundant calls and alleviate the impact. Around 16:20 UTC, we observed another performance degradation on QStash requests due to a load increase, and disabled event log
Timeline
- Postmortem · Apr 3, 15:49 UTC
### Product: QStash ### Incident Summary Due to high load, the volume of QStash event logs reached to a point which caused latency in the underlying data store operations. Event log creation was slowed down and lead to performance degradation in QStash request processing. In order to resolve the performance degradation in QStash requests, event logging module was turned off temporarily. After deploying a hot fix and configuration changes, we eventually turned on event logging and system went back to stable state again. ### Root Cause At 07:15 UTC we received alerts on the performance degradation and started the investigation. We discovered long running queries for synching event logs from main QStash servers to QStash event server. In order to resolve performance degradation, we turned off event logging functionality as an immediate action. This action turned the performance back to normal levels for QStash requests but left event log processing disabled. We deployed a hotfix during the day to remove some redundant calls and alleviate the impact. Around 16:20 UTC, we observed another performance degradation on QStash requests due to a load increase, and disabled event log processing again. In the following hours, we deployed a configuration change to relax the job interval durations for event log tasks and turned on event logging again. This configuration change helped to resolve the performance issues without any further issues. ### Impact During the problematic t
- Resolved · Apr 2, 21:12 UTC
QStash service and Event logs are fully functional without any remaning issues.
- Monitoring · Apr 2, 20:20 UTC
Monitoring: Main QStash service is back to normal. Event logs service is back online but events will be lagging a few mins.
- Monitoring · Apr 2, 17:33 UTC
Main QStash service is back to normal. Event logs are still temporarily unavailable.
- Investigating · Apr 2, 16:33 UTC
We are continuing to investigate this issue.
- Investigating · Apr 2, 16:29 UTC
We are continuing to investigate this issue.
- Investigating · Apr 2, 16:21 UTC
We are currently investigating this issue.
- Monitoring · Apr 2, 08:15 UTC
A fix has been implemented and we are monitoring the results.
- Investigating · Apr 2, 07:15 UTC
We are currently investigating this issue.
More from Upstash
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 913:00 UTC | Intermittent DNS Resolution Errors for Upstash Vector in US East (us-east-1) | minor | 0m |
| Sep 308:29 UTC | Upstash Console login issue | none | 0m |
| Aug 2822:56 UTC | Message Persistence Issue , QStash us-east-1 | major | 0m |
| Jul 2208:50 UTC | Fly.io infrastructure disruption affecting some Upstash Redis databases on Fly.io DFW Region | none | 1h 15m |
| Jul 1608:35 UTC | QStash us-east-1 - URL Publish Errors | none | 0m |
| Jun 2515:23 UTC | Upstash Redis Partial Service Disruption | major | 1h 12m |
Also caused by capacity and load
All| Started | Vendor | Incident | Impact | Duration |
|---|---|---|---|---|
| Sep 2213:19 UTC | We are investigating an issue with CH servers in Azure germanywestcentral region | minor | 5h 6m | |
| Sep 1607:20 UTC | Degradation with Gemini 3.8 Flash | major | 10h 28m | |
| Sep 1509:47 UTC | Disruption with some GitHub services | minor | 1h 30m | |
| Sep 422:02 UTC | Degradation in repos contents API | minor | 21m | |
| Sep 221:44 UTC | Elevated Linux worker queue times | minor | 21m | |
| Sep 115:00 UTC | Delays in commit processing | minor | 1h 1m |
From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.