Yandex Cloud
Search
Discuss with expertTry it for free
  • Customer Stories
  • Documentation
  • Blog
  • All Services
    • Cloud Interconnect
    • Cloud Backup
    • Cloud Registry
    • Yandex AI Studio
    • Compute Cloud
    • Object Storage
    • Managed Service for Kubernetes®
    • Yandex BareMetal
    • Smart Web Security
    • Security Deck
    • Managed Service for PostgreSQL
    • Managed Service for ClickHouse®
    • Monium
    • Cloud CDN
    • Network Load Balancer
    • Virtual Private Cloud
    • Cloud DNS
    • Application Load Balancer
    • Yandex Cloud Video
    • Stackland
    • Yandex Cloud Router
    • Yandex Managed Service for Trino
    • Managed Service for MySQL®
    • Managed Service for Valkey™
    • Managed Service for Apache Spark™
    • Yandex StoreDoc
    • Managed Service for OpenSearch
    • Managed Service for Apache Kafka®
    • Data Transfer
    • Yandex MPP Analytics Engine for PostgreSQL
    • Yandex Managed Service for Apache Airflow®
    • Data Processing
    • Yandex MetaData Hub
    • Managed Service for YDB
    • Managed Service for Sharded PostgreSQL
    • Managed Service for YTsaurus
    • Yandex WebSQL
    • DataLens
    • Yandex Search API
    • SpeechSense
    • SpeechKit
    • DataSphere
    • Vision OCR
    • Translate
    • Yandex Identity Hub
    • Key Management Service
    • Certificate Manager
    • Yandex Lockbox
    • Audit Trails
    • SmartCaptcha
    • Cloud Desktop
    • Yandex SIEM
    • SourceCraft Code Assistant
    • Container Registry
    • Managed Service for GitLab
    • Managed Service for Prometheus®
    • Cloud Functions
    • API Gateway
    • Yandex Cloud Postbox
    • Message Queue
    • Serverless Integrations
    • IoT Core
    • Data Streams
    • Serverless Containers
    • Cloud Notification Service
    • Yandex Query
    • Identity and Access Management
    • Yandex Cloud Console
    • Resource Manager
    • Yandex Cloud Billing
    • Yandex Cloud Quota Manager
    • Cloud Apps
  • System Status
  • Marketplace
    • Featured
    • Infrastructure & Network
    • Data Platform
    • AI for business
    • Security
    • DevOps tools
    • Serverless
    • Monitoring & Resources
  • All Solutions
    • By industry
    • By use case
    • Economics and Pricing
    • Security
    • Technical Support
    • Start testing with double trial credits
    • Cloud credits to scale your IT product
    • Gateway to Russia
    • Cloud for Startups
    • Center for Technologies and Society
    • Yandex Cloud Partner program
    • Price calculator
    • Pricing plans
  • Customer Stories
  • Documentation
  • Blog
© 2026 Direct Cursus Technology L.L.C.
Yandex MPP Analytics for PostgreSQL
  • Getting started
    • Overview of Greenplum® and Apache Cloudberry™ DBMSs in Yandex MPP Analytics for PostgreSQL
    • Resource relationships
    • Host classes
    • High availability clusters
    • Calculating the cluster configuration
    • Networking in Yandex MPP Analytics for PostgreSQL
    • Quotas and limits
    • Backups
    • Resource groups
    • Data distribution
    • Users and roles
    • User authentication
    • Command center
    • Command center settings
    • External tables
    • Managing connections
    • Expanding a cluster
    • Maintenance
    • Table and system folder vacuuming
    • DBMS settings
  • Access management
  • Inspections and recommendations
  • Pricing policy
  • Terraform reference
  • Monitoring metrics
  • Audit Trails events
  • Public materials
  • Release notes

In this article:

  • Choosing a distribution policy
  • Distribution key
  • Choosing a distribution key
  1. Concepts
  2. Data distribution

Data distribution in Yandex MPP Analytics for PostgreSQL

Written by
Yandex Cloud
Updated at September 23, 2026
View in Markdown
  • Choosing a distribution policy
  • Distribution key
    • Choosing a distribution key

In Yandex MPP Analytics for PostgreSQL, all database tables are distributed row by row across the cluster's segment hosts. A data distribution policy is a rule that defines how rows are distributed across segments.

Yandex MPP Analytics for PostgreSQL supports the following distribution policies:

  • Hash distribution: Rows are distributed across segments based on a hash function applied to the values of one or more columns defined as the distribution key.
  • Random distribution: Rows are distributed across all segments in a round-robin fashion.
  • Replicated distribution: A table copy is stored on each segment within the cluster.

You can define a distribution policy for a new table, as well as modify the distribution policy and key for an existing table. If you do not specify a distribution policy when creating a table, hash distribution is used. The PRIMARY KEY is used as the distribution key. If no primary key is defined, the first eligible column in the table is selected. For more on how the distribution key is selected when no policy is specified, see this Greenplum® reference and this Apache Cloudberry™ guide.

Choosing a distribution policyChoosing a distribution policy

In Yandex MPP Analytics for PostgreSQL, query performance depends on how data is distributed across the cluster's segments. If data distribution is uneven, segments will process unequal amounts of data. This can lead to imbalanced workloads and memory shortages on individual segments. JOIN operations also affect query performance. If tables use the same distribution key, rows with identical key values are stored on the same segment. This allows join operations to run locally, without transferring data between segments.

To distribute workloads evenly across segments, use the following rules when choosing a distribution policy:

  • Hash distribution: choose this policy if you can find a valid distribution key for the table.

  • Random distribution: use this strategy if no valid distribution key is available.

    With random distribution, individual INSERT and COPY operations may lead to data skew across segments. Additionally, tables with random distribution do not support local joins.

  • Replicated distribution: use for small tables, such as dictionaries.

Distribution keyDistribution key

A distribution key consists of one or more table columns whose values are used to assign rows to segments. Your choice of the distribution key controls how evenly your data is distributed across segments.

Choosing a distribution keyChoosing a distribution key

Select columns that are frequently used in JOIN operations as your distribution key. If the table has a primary key, the latter must include all columns of the distribution key. Similarly, if the table has a unique constraint, it must include all columns of the distribution key.

When choosing a distribution key, avoid columns that may lead to uneven data distribution:

  • Date and time columns.
  • Columns with multiple identical values.
  • Columns with a large number of NULL values.
  • Columns used in WHERE clauses.

Columns with geometric and custom data types cannot be used as a distribution key.

If a single column is not enough to distribute your data evenly, use a composite key consisting of two columns. Using a key of three or more columns does not improve data distribution balance, but only increases hashing time.

For more information, see this Greenplum® guide and this Apache Cloudberry™ article.

Greenplum® and Greenplum Database® are registered trademarks or trademarks of Broadcom Inc. in the United States and/or other countries.

Apache® and Apache Cloudberry™ are registered trademarks or trademarks of the Apache Software Foundation in the United States and/or other countries.

Was the article helpful?

Previous
Resource groups
Next
Users and roles
© 2026 Direct Cursus Technology L.L.C.