Yandex Cloud
Search
Discuss with expertTry it for free
  • Customer Stories
  • Documentation
  • Blog
  • All Services
    • Cloud Interconnect
    • Cloud Backup
    • Cloud Registry
    • Yandex AI Studio
    • Compute Cloud
    • Object Storage
    • Managed Service for Kubernetes®
    • Yandex BareMetal
    • Smart Web Security
    • Security Deck
    • Managed Service for PostgreSQL
    • Managed Service for ClickHouse®
    • Monium
    • Cloud CDN
    • Network Load Balancer
    • Virtual Private Cloud
    • Cloud DNS
    • Application Load Balancer
    • Yandex Cloud Video
    • Stackland
    • Yandex Cloud Router
    • Yandex Managed Service for Trino
    • Managed Service for MySQL®
    • Managed Service for Valkey™
    • Managed Service for Apache Spark™
    • Yandex StoreDoc
    • Managed Service for OpenSearch
    • Managed Service for Apache Kafka®
    • Data Transfer
    • Yandex MPP Analytics Engine for PostgreSQL
    • Yandex Managed Service for Apache Airflow®
    • Data Processing
    • Yandex MetaData Hub
    • Managed Service for YDB
    • Managed Service for Sharded PostgreSQL
    • Managed Service for YTsaurus
    • Yandex WebSQL
    • DataLens
    • Yandex Search API
    • SpeechSense
    • SpeechKit
    • DataSphere
    • Vision OCR
    • Translate
    • Yandex Identity Hub
    • Key Management Service
    • Certificate Manager
    • Yandex Lockbox
    • Audit Trails
    • SmartCaptcha
    • Cloud Desktop
    • SourceCraft Code Assistant
    • Container Registry
    • Managed Service for GitLab
    • Managed Service for Prometheus®
    • Cloud Functions
    • API Gateway
    • Yandex Cloud Postbox
    • Message Queue
    • Serverless Integrations
    • IoT Core
    • Data Streams
    • Serverless Containers
    • Cloud Notification Service
    • Yandex Query
    • Identity and Access Management
    • Yandex Cloud Console
    • Resource Manager
    • Yandex Cloud Billing
    • Yandex Cloud Quota Manager
    • Cloud Apps
  • System Status
  • Marketplace
    • Featured
    • Infrastructure & Network
    • Data Platform
    • AI for business
    • Security
    • DevOps tools
    • Serverless
    • Monitoring & Resources
  • All Solutions
    • By industry
    • By use case
    • Economics and Pricing
    • Security
    • Technical Support
    • Start testing with double trial credits
    • Cloud credits to scale your IT product
    • Gateway to Russia
    • Cloud for Startups
    • Center for Technologies and Society
    • Yandex Cloud Partner program
    • Price calculator
    • Pricing plans
  • Customer Stories
  • Documentation
  • Blog
© 2026 Direct Cursus Technology L.L.C.
Yandex Data Processing
  • Getting started
    • Resource relationships
    • Runtime environment
    • Yandex Data Processing component interfaces and ports
    • Yandex Data Processing jobs
    • Spark jobs
    • Autoscaling
    • Decommissioning of subclusters and hosts
    • Networking in Yandex Data Processing
    • Maintenance
    • Zones of control in Yandex Data Processing
    • Quotas and limits
    • Storage in Yandex Data Processing
    • Component properties
    • Apache Iceberg™ in Yandex Data Processing
    • Delta Lake in Yandex Data Processing
    • Yandex Data Processing logs
    • Initialization scripts
  • Access management
  • Pricing policy
  • Terraform reference
  • Monitoring metrics
  • Audit Trails events
  • Public materials
  • FAQ
  1. Concepts
  2. Apache Iceberg™ in Yandex Data Processing

Apache Iceberg™ in Yandex Data Processing

Written by
Yandex Cloud
Updated at June 29, 2026
View in Markdown

Apache Iceberg™ is an open table format for storing and processing large data arrays. It expands the feature set of the Apache Spark™ platform:

  • Supports the high-performance Apache Iceberg™ tables you use the same way as regular SQL tables.

  • Provides the schema evolution mechanism which eliminates side effects when updating schemas.

  • Provides hidden partitioning in auto mode thus preventing errors related to manual partitioning.

  • Allows retrospective requests enabled by the time travel mechanism. You can use the feature to make reproducible requests based on table snapshots or compare changes.

    Note

    This mechanism requires Apache Spark™ 3.3.x or higher.

  • Allows rolling tables back to previous versions (version rollback) for quick response to issues.

  • Provides advanced filtering that relies on column-level and partition-level statistics as well as table metadata. This accelerates request processing, even for very large tables: data files unrelated to the request will not be processed.

  • Enables the serializable isolation level — the strictest one for transaction isolation. All changes in tables are atomic, and the readers will see only the committed ones.

  • Supports concurrent writing based on the optimistic strategy: a writer will retry an operation if their changes are in conflict with those of another writer.

You can configure Apache Iceberg™ in a Yandex Data Processing cluster versions 2.0 or higher.

Note

Apache Iceberg™ is not part of Yandex Data Processing. It is not covered by Yandex Cloud support and its usage is not governed by the Yandex Data Processing Terms of Use.

For more information about Apache Iceberg™, see this official guide.

Compatibility between Apache Iceberg™ versions and Yandex Data Processing imagesCompatibility between Apache Iceberg™ versions and Yandex Data Processing images

Apache Iceberg™ versions and Yandex Data Processing images are only compatible if the Apache Iceberg™ version is compatible with the Apache Spark™ version used in the cluster. The table below lists compatible versions and links to the library files you will need to configure Apache Iceberg™ in your cluster.

Yandex Data Processing image

Apache Spark™ version

Apache Iceberg™ version

JAR files

2.0.x

3.0.3

1.0.0

iceberg-spark-runtime-3.0_2.12-1.0.0.jar

2.1.x

3.3.2

1.5.2

iceberg-spark-runtime-3.3_2.12-1.5.2.jar

2.2.x

3.5.0

1.5.2

iceberg-spark-runtime-3.5_2.12-1.5.2.jar

Note

Access to image 2.2 is provided on request. Contact support or your account manager.

Was the article helpful?

Previous
Component properties
Next
Delta Lake in Yandex Data Processing
© 2026 Direct Cursus Technology L.L.C.