We are investigating an issue with CH servers in Azure germanywestcentral region
Final update
The underlying AZ has recovered and we are fully back online. Since ClickHouse cloud architecture uses 3 AZs we were operational during this outage but there might have been some degradations. Azure support shared the details of the issue as following- Between 12:04 UTC and 17:46 UTC on 22 September 2022, a platform issue resulted in an impact to the Virtual Machines (VMs) and Virtual Machine Scale Sets (VMSS) in Germany West Central. Customers may have experienced their virtual machines becoming unresponsive, restarting unexpectedly, or being temporarily unavailable during the impact window.
Timeline
- Resolved · Sep 22, 18:25 UTC
The underlying AZ has recovered and we are fully back online. Since ClickHouse cloud architecture uses 3 AZs we were operational during this outage but there might have been some degradations. Azure support shared the details of the issue as following- Between 12:04 UTC and 17:46 UTC on 22 September 2022, a platform issue resulted in an impact to the Virtual Machines (VMs) and Virtual Machine Scale Sets (VMSS) in Germany West Central. Customers may have experienced their virtual machines becoming unresponsive, restarting unexpectedly, or being temporarily unavailable during the impact window.
- Monitoring · Sep 22, 17:02 UTC
Azure support shared they are working on fixing the underlying issue, we started seeing some recovery but the incident is not completely resolved yet. We are working with Azure to get this fixed as soon as possible.
- Monitoring · Sep 22, 15:26 UTC
The issue is still present on Azure infrastructure side and ClickHouse Cloud team is actively monitoring for any impact for existing ClickHouse Clusters. Correction on earlier impact description: Provisioning of new instances is expected to work. Operations like autoscaling of the ClickHouse Keeper replicas can be stuck or slow. Latency for insert and DDL queries are not affected, thanks to cross-AZ redundancy.
- Identified · Sep 22, 14:57 UTC
The Azure team is working on resolution of the issue. The ClickHouse Cloud team will relay any updates related to ClickHouse clusters operation. An excerpt from Azure incident status: > As per our investigation, we believe this stemmed from multiple underlying storage systems becoming unavailable at the same time, which led to attached virtual machines becoming unresponsive, restarting unexpectedly, or being temporarily unavailable.
- Investigating · Sep 22, 13:49 UTC
The problem is caused by Azure outage in the germanywestcentral region. ClickHouse Cloud team is working with Azure support. Impact: • New ClickHouse services can't be provisioned (appear as stuck in UI). • Some services may experience higher latency due to reduced capacity.
- Investigating · Sep 22, 13:31 UTC
Only one Availability zone is affected in germanywestcentral Azure region. The problem affects internal component (ClickHouse Keepers). Some ClickHouse Services in this region may see increased latency on inserts or DDL queries.
- Investigating · Sep 22, 13:21 UTC
Only germanywestcentral Azure region is affected.
- Investigating · Sep 22, 13:19 UTC
We detected an issue with CHC clusters in some Azure regions. ClickHouse cloud team is investigating. Some operations might be degraded.
More from ClickHouse
Full history| Started | Incident | Impact | Duration |
|---|---|---|---|
| Sep 2315:28 UTC | We are investigating an issue affecting Azure eastus2 region | minor | 30m |
| Sep 1610:16 UTC | Intermittent Issues with the Support tab | minor | 13h 50m |
| Sep 1518:57 UTC | ClickPipes failing on Kinesis in AWS us-east-1 | critical | 29h 46m |
| Sep 1420:36 UTC | Distributed Cache performance degradation in AWS us-east-1 | minor | 3h 59m |
| Sep 116:11 UTC | Service degradation in GCP us-central-1 | none | 19m |
| Aug 2310:14 UTC | Slow provisioning times of services in GCP europe-west4 region | minor | 32m |
Also caused by capacity and load
All| Started | Vendor | Incident | Impact | Duration |
|---|---|---|---|---|
| Sep 1607:20 UTC | Degradation with Gemini 3.8 Flash | major | 10h 28m | |
| Sep 1509:47 UTC | Disruption with some GitHub services | minor | 1h 30m | |
| Sep 422:02 UTC | Degradation in repos contents API | minor | 21m | |
| Sep 221:44 UTC | Elevated Linux worker queue times | minor | 21m | |
| Sep 115:00 UTC | Delays in commit processing | minor | 1h 1m | |
| Aug 2615:11 UTC | Incident with Actions | critical | 2h 50m |
From vendors' own status pages and disclosures. Times as reported. Logos via logo.dev; trademarks belong to their owners.