Cluster state monitoring Apache Hive™ Metastore
Data on the cluster and host state is available in the management console
Diagnostic information about cluster states is presented as graphs.
Charts are updated every 15 seconds.
Note
The most appropriate multiple units (MB, GB, and more) are automatically used in charts.
You can configure alerts in Yandex Monitoring to receive notifications about cluster failures. In Yandex Monitoring, there are two trigger thresholds: Warning and Alarm. If the specified threshold is exceeded, you will get alerts via the configured notification channels.
Cluster state monitoring
To view detailed information on the health state of a Apache Hive™ Metastore cluster:
-
In the management console
, select the folder. -
Navigate
to Yandex MetaData Hub. -
In the left-hand panel, select
Metastore. -
Click the cluster name and open the Monitoring tab.
-
To get started with Yandex Monitoring metrics, dashboards, or alerts, click Open in Monium in the top panel.
The page displays the following charts:
-
Under API:
- API Calls (rps): Methods with the largest number of executed requests for each cluster instance. The list of methods is updated dynamically. The chart may display up to ten methods.
- API Latency (p95): 95th percentile of the total request execution time broken down by instances.
- Active Calls: Number of requests being processed.
- API Latency by method (p95): Methods with the largest 95th percentile of request execution time for each cluster instance. The list of methods is updated dynamically. The chart may display up to ten methods.
-
Under Transactions & Errors:
- Open Transactions / Connections: Number of active transactions and connections:
- Open Transactions: Number of active ACID
transactions. - Open Connections: Number of active connections to the Apache Hive™ Metastore cluster over the Thrift protocol.
- Open Transactions: Number of active ACID
- Compaction Cycle Duration: Compression cycle duration.
- DirectSQL Errors: Number of errors during DirectSQL queries after which Apache Hive™ Metastore proceeded to execute queries via JDO.
- Open Transactions / Connections: Number of active transactions and connections:
-
Under Connection Pools (only for Apache Hive™ Metastore version
4.2):-
Connection States: Number of JDBC connections:
- Active Connections: Active connections.
- Idle Connections: Idle connections.
- Pending Connections: Number of queries awaiting connection allocation.
- Total Connections: Total number of connections.
-
Pool Usage (p95): 95th connection hold-time percentile of applications.
-
Connection Creation (p95): 95th connection setup time percentile.
-
Pool Connection Timeouts (rps): Number of queries per second that timed out due to absence of available connections.
-
-
Under DB Objects:
- Object Counts: Number of databases, tables, and partitions in the Apache Hive™ Metastore cluster storage.
- Object Operations (rps): Number of create and delete operations for databases, tables, and partitions, per second.
-
Under Resources:
- Instances: Number of instances per cluster.
- CPU Usage: CPU time usage by each instance as a percentage of total available CPU time.
- Memory Usage: RAM usage by each instance as a percentage of total available RAM.
-
Under JVM Memory:
- JVM Heap: Maximum amount of JVM heap memory and its usage by each instance.
- JVM Memory Pools: Pools with the largest amount of JVM memory used by each instance. The chart displays up to six pools. The list of pools is formed dynamically.
-
Under JVM Runtime:
- GC Rate: Operating speed of G1_Old_Generator and G1_Young_Generation garbage collectors on each instance.
- JVM Pauses: Number of stops of all JVM threads whose duration exceeded the notification and warning thresholds.
Setting up alerts in Yandex Monitoring
To configure cluster state indicator alerts:
- In the management console
, select the folder with the cluster for which you want to set up alerts. - Go
to Monitoring. - Under Service dashboards, select Managed Service For Hive Metastore – Cluster Overview.
- Click
on the relevant chart and select Create alert. - If the chart displays multiple metrics, select the data query for the relevant metric and click Continue. Learn more about the query language in this Yandex Monitoring guide.
- Set the
AlarmandWarningthreshold values to trigger the alert. - Click Create alert.
To have other cluster health indicators monitored automatically:
- Create an alert.
- Add a status metric.
- In the alert parameters, set the alert thresholds.
Recommended threshold values for selected metrics:
| Metric | Designation | Alarm |
Warning |
|---|---|---|---|
| CPU usage | cpu_usage |
90% | 80% |
| RAM usage | memory_usage |
90% | 80% |
| 95th percentile of total request execution time | metastore_api_calls_duration_seconds |
10 seconds | 3 seconds |
| Number of queries awaiting connection allocation | metastore_pool_pending_connections |
— | Above zero for five minutes |
| Number of queries that timed out due to absence of available connections | metastore_pool_connection_timeouts_total |
Greater than zero | — |
| Number of errors during DirectSQL queries after which Apache Hive™ Metastore proceeded to execute queries via JDO | metastore_directsql_errors_total |
— | Greater than zero |
For a full list of supported metrics, see this Monitoring guide.
Cluster health and status
A cluster’s State indicates its health, while its Status shows whether the cluster is started, stopped, or in a transitory state.
To view the health state and status of a cluster:
- In the management console
, select the folder. - Navigate
to Yandex MetaData Hub. - In the left-hand panel, select
Metastore. - In the cluster row, hover over the indicator in the Availability column.
Cluster health states
|
State |
Description |
Suggested actions |
|
ALIVE |
The cluster is operating normally. |
No action is required. |
|
DEGRADED |
The cluster is not running at its full capacity. |
Contact support
|
|
DEAD |
The cluster is out of order. |
Contact support
|
|
UNKNOWN |
The cluster’s state is unknown. |
Contact support
|
Cluster statuses
| Status | Description | Suggested actions |
|---|---|---|
| CREATING | Preparing for the first start | Wait a while and get started. The time it takes to create a cluster depends on the host class. |
| RUNNING | The cluster is operating normally | No action is required. |
| STOPPING | The cluster is stopping | After a while, the cluster status will switch to STOPPED and the cluster will be disabled. No action is required. |
| STOPPED | The cluster is stopped | Start the cluster to get it running again. |
| STARTING | Starting the cluster that was stopped earlier | After a while, the cluster status will switch to RUNNING. Wait a while and get started. |
| UPDATING | Updating the cluster's configuration | Once the update is complete, the cluster will get the status it had prior to the update: RUNNING or STOPPED. |
| ERROR | Error when performing an operation with the cluster or during a maintenance window | If the cluster remains in this status for a long time, contact support |
| STATUS_UNKNOWN | The cluster is unable to determine its status | If the cluster remains in this status for a long time, contact support |