Skip to content
Upstash · Data and observabilityApr 2, 2025, 07:15 UTC

QStash: Degraded performance in request processing and event logs

MinorCapacityUpdated 20h ago
Apr 2, 07:15 UTCApr 2, 21:12 UTC
Duration
13h 57m
Impact
Minor
Root cause
Capacity
Upstash, 90 days
5 incidents
Affected
EU-CENTRAL-1QStash
Status page

Final update

### Product: QStash ### Incident Summary Due to high load, the volume of QStash event logs reached to a point which caused latency in the underlying data store operations. Event log creation was slowed down and lead to performance degradation in QStash request processing. In order to resolve the performance degradation in QStash requests, event logging module was turned off temporarily. After deploying a hot fix and configuration changes, we eventually turned on event logging and system went back to stable state again. ### Root Cause At 07:15 UTC we received alerts on the performance degradation and started the investigation. We discovered long running queries for synching event logs from main QStash servers to QStash event server. In order to resolve performance degradation, we turned off event logging functionality as an immediate action. This action turned the performance back to normal levels for QStash requests but left event log processing disabled. We deployed a hotfix during the day to remove some redundant calls and alleviate the impact. Around 16:20 UTC, we observed another performance degradation on QStash requests due to a load increase, and disabled event log

Timeline

  1. Postmortem · Apr 3, 15:49 UTC
    ### Product: QStash ### Incident Summary Due to high load, the volume of QStash event logs reached to a point which caused latency in the underlying data store operations. Event log creation was slowed down and lead to performance degradation in QStash request processing. In order to resolve the performance degradation in QStash requests, event logging module was turned off temporarily. After deploying a hot fix and configuration changes, we eventually turned on event logging and system went back to stable state again. ### Root Cause At 07:15 UTC we received alerts on the performance degradation and started the investigation. We discovered long running queries for synching event logs from main QStash servers to QStash event server. In order to resolve performance degradation, we turned off event logging functionality as an immediate action. This action turned the performance back to normal levels for QStash requests but left event log processing disabled. We deployed a hotfix during the day to remove some redundant calls and alleviate the impact. Around 16:20 UTC, we observed another performance degradation on QStash requests due to a load increase, and disabled event log processing again. In the following hours, we deployed a configuration change to relax the job interval durations for event log tasks and turned on event logging again. This configuration change helped to resolve the performance issues without any further issues. ### Impact During the problematic t
  2. Resolved · Apr 2, 21:12 UTC
    QStash service and Event logs are fully functional without any remaning issues.
  3. Monitoring · Apr 2, 20:20 UTC
    Monitoring: Main QStash service is back to normal. Event logs service is back online but events will be lagging a few mins.
  4. Monitoring · Apr 2, 17:33 UTC
    Main QStash service is back to normal. Event logs are still temporarily unavailable.
  5. Investigating · Apr 2, 16:33 UTC
    We are continuing to investigate this issue.
  6. Investigating · Apr 2, 16:29 UTC
    We are continuing to investigate this issue.
  7. Investigating · Apr 2, 16:21 UTC
    We are currently investigating this issue.
  8. Monitoring · Apr 2, 08:15 UTC
    A fix has been implemented and we are monitoring the results.
  9. Investigating · Apr 2, 07:15 UTC
    We are currently investigating this issue.

More from Upstash

Full history

Also caused by capacity and load

All
StartedIncidentDuration
Sep 2213:19 UTCWe are investigating an issue with CH servers in Azure germanywestcentral regionClickHouse5h 6m
Sep 1607:20 UTCDegradation with Gemini 3.8 FlashGitHub10h 28m
Sep 1509:47 UTCDisruption with some GitHub servicesGitHub1h 30m
Sep 422:02 UTCDegradation in repos contents APIGitHub21m
Sep 221:44 UTCElevated Linux worker queue timesExpo21m
Sep 115:00 UTCDelays in commit processingGitHub1h 1m

From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.

Weekly: the week's major outages, postmortems and breaches, Saturday mornings.