Inspections and recommendations in Managed Service for ClickHouse®
Managed database clusters regularly undergo diagnostics to detect possible issues, increase the cluster's reliability, and improve its performance. The results of such checks are displayed as inspections under Recommendations. You can see notifications about successful checks and recommendations on eliminating discovered risks. The responsibility to troubleshoot any detected issues lies within the Yandex Cloud user's remit.
All checks have a severity level:
- High level: Criteria of high cluster availability, significant risks of reduced performance, data loss risks. Such checks warrant special attention and require following the recommendations provided.
- Moderate level: Possible risks of reduced performance, suboptimal memory and disk space usage.
- Low level: Potential risks and cluster operation limitations.
The list of recommendations is updated regularly. For each recommendation, the following timestamps are fixed: the date of the risk's first detection and the date of the most recent status update on this issue. If the recommendation seems to be excessive or incorrect, you can hide it, specifying a reason. Once the hiding period expires, the recommendation will automatically become available again if the issue persists.
Recommendations are available at the cluster, folder, and cloud levels and provide tips for all your resources. However, the absence of recommendations does not mean that your cluster is optimized: the list of checks gets continuously appended but still remains incomprehensive and cannot replace monitoring, since it is targeted at detecting patterns rather than specific issues. You can additionally run cluster performance diagnostics and analyze monitoring metrics.
Managing recommendations in Managed Service for ClickHouse® requires the managed-clickhouse.editor role or higher.
Available inspections in Managed Service for ClickHouse®
| Category | Check | Risk | Severity |
|---|---|---|---|
| Performance | CPU usage | High CPU usage on the host | High |
| High availability | Allocating disk space | Insufficient disk space | High |
| High availability | Hosting the coordination service on separate hosts | Risk of unavailability due to hosting the coordination service on hosts | High |
| High availability | High shard availability | Risk of data loss and cluster unavailability if an availability zone fails | High |
| High availability | Coordination service quorum in case of a zone failure | Risk of losing the coordination service quorum in the event of a zone failure | High |
High CPU usage on the host
Description
The host continuously uses all CPU resources, which may cause delays in query processing. Check the load and optimize your queries (which may include using the WebSQL AI assistant) or increase the cluster's computing resources.
Action
To increase the cluster's computing resources:
- Navigate to Managed Service for ClickHouse.
- Select your cluster and click
Edit. - Under Resources, select a host class with the required amount of vCPUs.
- Click Save changes.
Insufficient disk space
Description
Your host is critically low on free space. The cluster may become unavailable. Increase the disk size or clear unused data.
Action
To increase the disk size:
- Navigate to Managed Service for ClickHouse.
- Select your cluster and click
Edit. - Under Storage size, increase the disk size.
- Click Save changes.
Risk of unavailability due to hosting the coordination service on hosts
Description
A configuration where ClickHouse® and ClickHouse® Keeper share the hosts is not highly available. Place ClickHouse® Keeper on separate hosts.
Note
Single-host clusters are excluded from high availability testing: Yandex Cloud users are fully responsible for managing such a configuration.
Action
To update the ClickHouse® Keeper settings:
- In the management console
, navigate to the folder dashboard and select Managed Service for ClickHouse. - Select your cluster and click Edit in the top panel.
- For the ClickHouse Keeper (on separate hosts) coordination service, select the platform, VM type, and host class under Clickhouse Keeper host class.
Risk of data loss and cluster unavailability if an availability zone fails
Description
- The shard has no replicas, which poses a risk of data loss and cluster unavailability. Make sure the specified shards have replicas in another zone.
- The current distribution of shard hosts does not ensure fault protection in a single availability zone. Make sure the replicas are distributed across different availability zones.
Action
To move ClickHouse® hosts:
-
In the management console
, select the folder containing the cluster. -
Navigate
to Managed Service for ClickHouse. -
Click the cluster name and navigate to the Hosts tab.
-
Click Create host.
-
Specify the following host settings:
- Target availability zone for your hosts.
- New subnet.
- To make the host accessible from outside Yandex Cloud, select Public access.
-
Click Save.
For more information, see this host migration guide.
Risk of losing the coordination service quorum in the event of a zone failure
Description
Your coordination service hosts are distributed unevenly: one of the availability zones has too few hosts to form a quorum. If that zone fails, the remaining hosts may lose quorum, blocking leader elections and stopping normal cluster operation. Distribute your hosts across zones so that losing any single zone does not lead to a loss of quorum.
Action
To move ZooKeeper hosts:
- Create a subnet in the target availability zone for the hosts.
- In the management console
, select the folder containing the cluster. - Navigate
to Managed Service for ClickHouse. - Click the cluster name and navigate to the Hosts tab.
- Click Set up coordinator service.
- Specify the new subnet and the availability zone to move the hosts to.
- Click Save.
For more information, see this host migration guide.