Yandex Cloud
Search
Discuss with expertTry it for free
  • Customer Stories
  • Documentation
  • Blog
  • All Services
    • Cloud Interconnect
    • Cloud Backup
    • Cloud Registry
    • Yandex AI Studio
    • Compute Cloud
    • Object Storage
    • Managed Service for Kubernetes®
    • Yandex BareMetal
    • Smart Web Security
    • Security Deck
    • Managed Service for PostgreSQL
    • Managed Service for ClickHouse®
    • Monium
    • Cloud CDN
    • Network Load Balancer
    • Virtual Private Cloud
    • Cloud DNS
    • Application Load Balancer
    • Yandex Cloud Video
    • Stackland
    • Yandex Cloud Router
    • Yandex Managed Service for Trino
    • Managed Service for MySQL®
    • Managed Service for Valkey™
    • Managed Service for Apache Spark™
    • Yandex StoreDoc
    • Managed Service for OpenSearch
    • Managed Service for Apache Kafka®
    • Data Transfer
    • Yandex MPP Analytics Engine for PostgreSQL
    • Yandex Managed Service for Apache Airflow®
    • Data Processing
    • Yandex MetaData Hub
    • Managed Service for YDB
    • Managed Service for Sharded PostgreSQL
    • Managed Service for YTsaurus
    • Yandex WebSQL
    • DataLens
    • Yandex Search API
    • SpeechSense
    • SpeechKit
    • DataSphere
    • Vision OCR
    • Translate
    • Yandex Identity Hub
    • Key Management Service
    • Certificate Manager
    • Yandex Lockbox
    • Audit Trails
    • SmartCaptcha
    • Cloud Desktop
    • Yandex SIEM
    • SourceCraft Code Assistant
    • Container Registry
    • Managed Service for GitLab
    • Managed Service for Prometheus®
    • Cloud Functions
    • API Gateway
    • Yandex Cloud Postbox
    • Message Queue
    • Serverless Integrations
    • IoT Core
    • Data Streams
    • Serverless Containers
    • Cloud Notification Service
    • Yandex Query
    • Identity and Access Management
    • Yandex Cloud Console
    • Resource Manager
    • Yandex Cloud Billing
    • Yandex Cloud Quota Manager
    • Cloud Apps
  • System Status
  • Marketplace
    • Featured
    • Infrastructure & Network
    • Data Platform
    • AI for business
    • Security
    • DevOps tools
    • Serverless
    • Monitoring & Resources
  • All Solutions
    • By industry
    • By use case
    • Economics and Pricing
    • Security
    • Technical Support
    • Start testing with double trial credits
    • Cloud credits to scale your IT product
    • Gateway to Russia
    • Cloud for Startups
    • Center for Technologies and Society
    • Yandex Cloud Partner program
    • Price calculator
    • Pricing plans
  • Customer Stories
  • Documentation
  • Blog
© 2026 Direct Cursus Technology L.L.C.
Yandex Data Processing
  • Getting started
    • Resource relationships
    • Runtime environment
    • Yandex Data Processing component interfaces and ports
    • Yandex Data Processing jobs
    • Spark jobs
    • Autoscaling
    • Decommissioning of subclusters and hosts
    • Networking in Yandex Data Processing
    • Maintenance
    • Zones of control in Yandex Data Processing
    • Quotas and limits
    • Storage in Yandex Data Processing
    • Component properties
    • Apache Iceberg™ in Yandex Data Processing
    • Delta Lake in Yandex Data Processing
    • Yandex Data Processing logs
    • Initialization scripts
  • Access management
  • Pricing policy
  • Terraform reference
  • Monitoring metrics
  • Audit Trails events
  • Public materials
  • FAQ
  1. Concepts
  2. Autoscaling

Autoscaling of subclusters

Written by
Yandex Cloud
Updated at July 22, 2026
View in Markdown

Note

Autoscaling of subclusters is supported for Yandex Data Processing clusters version 1.4 and higher.

Yandex Data Processing supports autoscaling of data processing subclusters based on metrics received by Yandex Monitoring:

  • If the metric value exceeds the specified threshold, new hosts will be added to a subcluster. You can start using them in a YARN cluster running Apache Spark or Apache Hive as soon as the host status changes to Alive.
  • If the value of the key metric falls below the specified threshold, the system will first start decommissioning and then removing redundant hosts in the subcluster.

You can read more about autoscaling in the Instance Groups documentation.

You can choose the scaling method that best suits your needs:

  • Default scaling: Scaling based on the yarn.cluster.containersPending metric.

    This internal YARN metric tracks the number of resource allocation units expected by jobs waiting in the queue. It is suitable for clusters with multiple jobs managed by Apache Hadoop® YARN. This scaling method does not require any additional configuration.

  • CPU utilization target, %: Scaling based on the vCPU usage metric. You can learn more about this type of scaling in the Instance Groups documentation.

To set up autoscaling of your cluster based on other metrics and formulas, contact support.

You can set the following autoscaling parameters:

  • Initial (minimum) size of the group.
  • Decommissioning timeout in seconds. The maximum value is 86400 seconds (24 hours). The default value is 120 seconds.
  • Type of VM instances: standard or preemptible.
  • Maximum group size.
  • Time period for calculating the average load on each VM instance in the group.
  • Instance warmup period: Interval during which instance metrics are not used after it starts. Average metric values for the group are used instead.
  • Stabilization period (minutes or seconds): Interval during which the number of instances in the group cannot be decreased.

Was the article helpful?

Previous
Spark jobs
Next
Decommissioning of subclusters and hosts
© 2026 Direct Cursus Technology L.L.C.