Yandex Cloud
Search
Discuss with expertTry it for free
  • Customer Stories
  • Documentation
  • Blog
  • All Services
    • Cloud Interconnect
    • Cloud Backup
    • Cloud Registry
    • Yandex AI Studio
    • Compute Cloud
    • Object Storage
    • Managed Service for Kubernetes®
    • Yandex BareMetal
    • Smart Web Security
    • Security Deck
    • Managed Service for PostgreSQL
    • Managed Service for ClickHouse®
    • Monium
    • Cloud CDN
    • Network Load Balancer
    • Virtual Private Cloud
    • Cloud DNS
    • Application Load Balancer
    • Yandex Cloud Video
    • Stackland
    • Yandex Cloud Router
    • Yandex Managed Service for Trino
    • Managed Service for MySQL®
    • Managed Service for Valkey™
    • Managed Service for Apache Spark™
    • Yandex StoreDoc
    • Managed Service for OpenSearch
    • Managed Service for Apache Kafka®
    • Data Transfer
    • Yandex MPP Analytics Engine for PostgreSQL
    • Yandex Managed Service for Apache Airflow®
    • Data Processing
    • Yandex MetaData Hub
    • Managed Service for YDB
    • Managed Service for Sharded PostgreSQL
    • Managed Service for YTsaurus
    • Yandex WebSQL
    • DataLens
    • Yandex Search API
    • SpeechSense
    • SpeechKit
    • DataSphere
    • Vision OCR
    • Translate
    • Yandex Identity Hub
    • Key Management Service
    • Certificate Manager
    • Yandex Lockbox
    • Audit Trails
    • SmartCaptcha
    • Cloud Desktop
    • SourceCraft Code Assistant
    • Container Registry
    • Managed Service for GitLab
    • Managed Service for Prometheus®
    • Cloud Functions
    • API Gateway
    • Yandex Cloud Postbox
    • Message Queue
    • Serverless Integrations
    • IoT Core
    • Data Streams
    • Serverless Containers
    • Cloud Notification Service
    • Yandex Query
    • Identity and Access Management
    • Yandex Cloud Console
    • Resource Manager
    • Yandex Cloud Billing
    • Yandex Cloud Quota Manager
    • Cloud Apps
  • System Status
  • Marketplace
    • Featured
    • Infrastructure & Network
    • Data Platform
    • AI for business
    • Security
    • DevOps tools
    • Serverless
    • Monitoring & Resources
  • All Solutions
    • By industry
    • By use case
    • Economics and Pricing
    • Security
    • Technical Support
    • Start testing with double trial credits
    • Cloud credits to scale your IT product
    • Gateway to Russia
    • Cloud for Startups
    • Center for Technologies and Society
    • Yandex Cloud Partner program
    • Price calculator
    • Pricing plans
  • Customer Stories
  • Documentation
  • Blog
© 2026 Direct Cursus Technology L.L.C.
Yandex DataSphere
  • Getting started
    • About Yandex DataSphere
    • DataSphere resource relationships
    • Communities
    • Cost management
    • Project
    • Computing resource configurations
      • Nodes and aliases
      • Health checks and monitoring
      • Node metrics
    • Foundation models
    • Quotas and limits
    • Special terms for educational institutions
  • Terraform reference
  • Audit Trails events
  • Access management
  • Pricing policy
  • Public materials
  • Release notes
  1. Concepts
  2. DataSphere Inference
  3. Health checks and monitoring

Health checks and monitoring

Written by
Yandex Cloud
Updated at March 4, 2024
View in Markdown

You can enable health checks for your node instances: the balancer will send check requests to endpoints at certain intervals and wait for a response for a certain period of time.

Checks can be implemented using HTTP or gRPC. The protocol must match the check implementation inside the node container.

The following health check settings are supported:

  • Timeout: Response waiting time.
  • Interval: Time interval between health check requests.
  • Resource health indicators: Successful or failed result thresholds. If a threshold is exceeded, the check passed or failed, respectively.
  • HTTP health check settings:
    • Path in the URI of request to the endpoint.
  • Settings of gRPC health checks:
    • Name of the service checked.

MonitoringMonitoring

Nodes supply monitoring metrics to the Yandex Monitoring service directory specified in the node settings. By default, the platform collects the following metrics:

  • For nodes:

    • node_requests: Frequency of requests to node, requests per second.
    • node_grpc_codes: Frequency of response codes for gRPC endpoints, codes per second for each code.
    • node_http_codes: Frequency of response codes for HTTP endpoints, codes per second for each code.
    • node_requests_durations: Request execution time histogram, in milliseconds.
  • For aliases:

    • alias_requests: Frequency of requests to an alias, requests per second.
    • alias_grpc_codes: Frequency of response codes for gRPC endpoints, codes per second for each code.
    • alias_http_codes: Frequency of response codes for HTTP endpoints, codes per second for each code.
    • alias_requests_durations: Request execution time histogram, in milliseconds.

Node and alias metrics contain additional labels:

  • node_id: Node ID
  • node_path: Path in the URI of request to the endpoint
  • alias_name: Alias name

You can get standard metrics using requests in Monitoring or from the DataSphere service dashboards on the node and alias pages.

Additionally, for nodes, you can enable export of any metrics to Monitoring. The platform will poll all node instances over HTTP and collect custom metrics every now and then. The charts will also be available in the Monitoring directory specified in the node settings.

The following settings are supported for collecting monitoring metrics:

  • Format: Prometheus text format or Monitoring format
  • HTTP path: GET request path
  • Port: Container port for HTTP requests

The following labels are automatically added to all metrics:

  • node_id: Node ID
  • instance_id: Node instance ID

Was the article helpful?

Previous
Nodes and aliases
Next
Node metrics
© 2026 Direct Cursus Technology L.L.C.