• novità

    • Release 2026.06 - Portiamo la data observability nel vostro codice

  • novità

    • Contribuite al futuro dell’innovazione in IA e dati

Iceberg Open Table Format Explained Simply

|

8

min di lettura

Your dashboard says yesterday's revenue dropped, but the problem is less obvious. One engine read a directory before a new batch was fully committed, another interpreted a partition differently, and a third still sees an older schema. The files exist, yet the table can't provide one dependable answer.

That's the problem the Iceberg open table format addresses. It adds a structured metadata and transaction layer to data stored in open files, while leaving teams free to use engines such as Spark, Flink, Trino, or Hive. The important decision in 2026, however, isn't only whether Iceberg is a good table format. It's which catalog controls the metadata, policies, and access path across those engines.

A digital illustration of an iceberg representing data hidden beneath the surface with servers and documents.

Table of Contents

Introduction to Modern Table Formats and Why They Matter

A traditional data lake often begins with an ingestion job writing Parquet files to object storage, a metastore recording a table location, and a query engine discovering files by scanning folders. That approach can work for append-only data, but operational pressure arrives quickly.

A producer may write into the wrong date directory. A repair job may alter files while a dashboard query is running. An analyst may read a mixture of old and new files. If the table's partition layout changes, every producer and consumer may need to understand the change manually. The data lake then behaves less like a database and more like a shared folder with undocumented rules.

Open table formats improve that arrangement by placing a table abstraction above the files. They track schemas, snapshots, file-level metadata, and commits so query engines can understand which files belong to a consistent version of the table. The data remains in open storage formats, but users gain database-like behavior.

For digna customers and similar enterprise teams, reliability means more than successful queries. A reporting table must arrive on time, retain the expected structure, pass business validations, and expose enough history to investigate an incident. Iceberg helps establish the table's transactional and structural foundation. Observability still needs to operate around that foundation, checking whether data arrived, whether values changed unexpectedly, and whether downstream consumers remain compatible. Teams exploring the wider concept can use this open table format overview as additional context.

Practical rule: Treat an open table format as a reliability layer for files, not as a complete governance platform.

Apache Iceberg was created at Netflix in 2017 to address scalability and consistency problems in Hive-style tables. Netflix donated it to the Apache Software Foundation in November 2018, and the project graduated to top-level Apache status in May 2020. Its specification continued through the stable 1.0.0 release in October 2022 and the 1.6.1 release in August 2024, according to the Apache Iceberg project history.

This guide starts with the basic table model, then moves through Iceberg's metadata architecture, transactional behavior, format comparisons, migration patterns, and production operations. The final focus is governance, because the file specification is only one part of an enterprise data platform.

What an Open Table Format Is and How It Works

Think of a raw data lake as a warehouse filled with labeled boxes. The Parquet files are the boxes, folders are rough storage areas, and a query engine is a worker searching for the boxes that might contain an answer. Without a reliable inventory, the worker may search too much, miss relevant files, or combine boxes from incompatible deliveries.

An open table format adds that inventory and a set of rules. It usually involves three connected layers:

  1. Data files hold the actual records, commonly in columnar formats such as Parquet.

  2. Table metadata describes schemas, snapshots, files, statistics, and partition information.

  3. A catalog gives engines a stable way to find the current table metadata and apply access policies.

That third layer is easy to underestimate. The format defines how table state is represented, while the catalog defines how engines locate that state. A Spark job, a Flink pipeline, and a Trino query can work with the same table only if they can resolve the table through a compatible catalog and interpret the relevant format version.

A diagram illustrating the four key components of an open table format for a data lake.

From files to a managed table

Suppose a pipeline receives sales records. It writes Parquet data files to object storage, but it doesn't place them in a date folder and hope that every reader understands the layout. Iceberg records the files in manifests, associates them with a table snapshot, and publishes a new metadata state through the catalog.

A reader first discovers the current table state. It then uses metadata to plan the scan, selecting only files that might contain matching records. This is different from asking the engine to list every object and inspect each file.

The word open refers to more than open-source code. An open table format aims to provide a documented specification that multiple engines can implement. Iceberg's specification has progressed through formal versions. The Apache documentation states that versions 1, 2, and 3 are complete and community-adopted, while version 4 remains under active development. The documentation also lists 1.11.0 as the latest documentation version in 2026, reflecting continued project investment in the Iceberg specification and documentation.

Openness reduces dependence on one query engine, but it doesn't automatically eliminate vendor-specific behavior. Readers and writers still need compatible feature support, and the catalog may apply its own policies, credentials, or governance rules. That distinction becomes central once several engines share production tables.

Inside the Iceberg Open Table Format Architecture

Iceberg's architecture is easiest to understand from the catalog downward. The catalog points to the current metadata file. That metadata identifies the active snapshot and connects the table to one or more manifest lists. Manifest lists point to manifest files, and manifests describe the data files available to the snapshot.

A diagram illustrating the Iceberg architecture showing the hierarchical relationship between catalogs, metadata, manifests, and data files.

Following a table read

A simplified read looks like this:

  1. The engine asks the catalog for the table's current metadata location.

  2. The metadata file identifies the current snapshot.

  3. The snapshot references a manifest list.

  4. The manifest list identifies manifest files.

  5. The engine evaluates manifest entries and reads eligible data files.

The data itself remains in files such as Parquet. For readers who want a separate explanation of the underlying storage format, this Parquet guide provides useful background. Iceberg's contribution is the table-level organization around those files.

Manifests store partition values for data files. When a query includes a predicate, the engine can compare that predicate with the partition tuples and reject files that can't match. This is metadata-driven pruning. The engine spends less effort opening irrelevant files, and the planning process doesn't depend on manually interpreting directory names.

Iceberg also supports hidden partitioning. A producer can write a timestamp while the table applies a partition transform internally, such as extracting a date or truncating a value. Producers and consumers don't need to manage a separate partition column directly, which reduces errors caused by inconsistent partition expressions.

Why snapshots matter

Every write creates a new snapshot. A reader resolves one snapshot and plans against that consistent state, rather than observing a table while files are being added or removed. This design supports snapshot isolation and serializable, atomic table changes, so readers don't see partial or uncommitted writes.

The snapshot model also enables time travel. An analyst investigating a dashboard discrepancy can query an earlier table state, provided the relevant snapshots and files remain available. That makes historical investigation more precise than relying on copied extracts or manually preserved folders.

Apache Iceberg's performance documentation notes that the metadata architecture can allow even multi-petabyte tables to be read from a single node without requiring a distributed SQL engine merely to sift through table metadata. The same documentation cites cases producing a 10x performance improvement, as described in the Iceberg performance documentation. The practical lesson is that metadata design can influence planning cost as much as file layout influences scan cost.

Transactional Guarantees Schema Evolution and Partitioning Strategies

Iceberg's value becomes clearer when a table changes while people and pipelines continue using it. A writer doesn't publish a half-finished table state. It prepares a new snapshot and commits that state atomically. Readers continue seeing the previous committed snapshot until the new one becomes visible.

This gives teams snapshot isolation and serializable table changes. Concurrent writers still need a conflict strategy, and a failed commit may require retry logic, but readers won't accidentally combine an incomplete write with an older table state.

A diagram outlining the key architectural guarantees of Apache Iceberg, including transactional guarantees, schema evolution, and partition evolution.

Changes that avoid data rewrites

Iceberg supports schema evolution operations such as adding, dropping, renaming, and reordering columns without rewriting existing data files. The metadata tracks the logical schema, so a column rename doesn't have to mean physically rewriting every historical file.

That capability doesn't remove the need for discipline. A renamed field may still confuse downstream tools that identify columns by name, and a dropped field may affect reports or models. Schema compatibility checks and consumer notifications remain important, especially when different engines support format features at different levels. A schema tracking system such as digna Schema Tracker can sit alongside the table and identify structural changes before they become downstream incidents.

Partition evolution follows a similar principle. Iceberg treats a partition-spec change as a metadata operation. Old files remain in their existing layout, while new files use the updated specification. Teams can adapt partitioning to new access patterns without rewriting the entire historical table or taking it offline.

Choosing hidden partitioning thoughtfully

Hidden partitioning is useful when teams want physical organization without exposing partition mechanics to every producer. A timestamp field can support date-based pruning without requiring every ingestion job to populate a matching partition column correctly.

There is a tradeoff. After partition evolution, a table may contain files written under different specifications. Iceberg can plan across that history, but operators still need to monitor whether old and new layouts produce uneven scan behavior. Partition changes should follow observed query patterns, not become a substitute for understanding workload access.

Row-level deletes

Iceberg format v2 supports row-level deletes through separate delete files rather than rewriting complete data files. Positional deletes identify a row by its file path and row position. Equality deletes identify rows by matching column values, which can support update and delete workflows with lower write amplification, as described in this Iceberg row-level delete explanation.

Delete files introduce maintenance considerations. Queries may need to merge base data and delete information, and compaction or rewrite operations may later consolidate the layout. The format reduces the immediate rewrite burden, but it doesn't eliminate lifecycle management.

Comparing Iceberg Delta Lake and Hudi for Your Lakehouse

Iceberg, Delta Lake, and Hudi all address the gap between unmanaged data files and database-like table behavior. The right choice depends on the workload, engines, catalog strategy, and operational services your team is prepared to run.

Criteria

Apache Iceberg

Delta Lake

Apache Hudi

Primary fit

Multi-engine analytic tables and open interoperability

Strong alignment with the Databricks ecosystem

Update-heavy and streaming-oriented lakehouse workloads

Metadata model

Snapshot, manifest-list, and manifest architecture

Transaction log-centered design

Timeline and table-service-oriented design

Schema evolution

Supports changes such as add, drop, rename, and reorder

Supports schema evolution, with behavior depending on runtime and feature support

Supports broad schema evolution capabilities

Partition strategy

Hidden partitioning and metadata-only partition evolution

Partition management depends on table and platform configuration

Uses partitioning plus clustering and table services

Row changes

v2 delete files, with positional and equality deletes

Updates and deletes through Delta operations

Strong support for upserts, deletes, and incremental workflows

Engine choice

Designed for broad engine interoperability

Particularly natural inside Databricks deployments

Strong Spark and streaming integration, with wider ecosystem support

Catalog question

Catalog is central to discovery and governance

Catalog and platform services strongly affect openness

Catalog may be less central to basic table operation, but governance still needs surrounding services

These descriptions are architectural guidance, not a universal benchmark. Query speed depends on file sizes, data distribution, sort order, compaction, engine versions, workload shape, and catalog implementation. A platform team should test representative reads and writes rather than choosing from feature checkboxes alone.

Match the format to the operating model

Choose Iceberg when several engines must share tables and the team values a broadly specified table layer with hidden partitioning and independent evolution. Choose Delta Lake when deep Databricks integration, native platform services, and a unified vendor operating model outweigh the need for broader catalog neutrality.

Hudi deserves consideration when frequent updates, change capture, incremental processing, or streaming ingestion drive the design. Its approach includes storage-engine and table-management capabilities that can shape how the platform handles indexing, compaction, and ingestion.

The more important question is often outside the file specification. A governance team needs one place to manage namespaces, ownership, classification, permissions, and audit history. Iceberg supplies table transactions and metadata, but it doesn't by itself provide a complete cross-engine policy system. Teams evaluating Iceberg alongside Databricks workloads can also review Databricks data quality considerations, particularly where platform controls and data reliability processes meet.

Ecosystem Integrations Migration Patterns and Example Workflows

Iceberg fits into a lakehouse through two interfaces, storage and catalog. Spark or Flink may produce the data, Trino or Hive may query it, and object storage may hold the files. The catalog ties those activities to one table identity and exposes the metadata needed by each compatible engine.

A diagram outlining the five steps of working with Apache Iceberg: Ingest, Catalog, Write, Query, and Evolve.

A representative workflow looks like this:

  • Ingest: Spark or Flink receives batch or streaming records.

  • Catalog: The table is registered through a catalog such as a Hive Metastore, AWS Glue Catalog, or an Iceberg REST catalog.

  • Write: The engine creates data files and publishes an atomic snapshot.

  • Query: Trino, Hive, Spark, or another compatible engine resolves the table through the catalog.

  • Evolve: The owner changes the schema or partition specification as the workload develops.

The exact syntax varies by engine, but the concepts remain stable. A SQL workflow might create a table with an explicit schema, insert records, alter a column, and query a prior snapshot. A Spark job might use Iceberg's DataFrameWriterV2 interface, while Flink uses its Iceberg sink and catalog configuration.

Migration without losing control

A Hive-to-Iceberg migration can follow different paths. Teams may convert existing data and register it through Iceberg metadata, or they may rewrite files to improve layout, normalize schemas, and remove legacy partition assumptions. Metadata conversion can reduce disruption, while a rewrite creates an opportunity to optimize data files. The choice depends on the condition of the existing table and the tolerance for migration work.

A Delta-to-Iceberg migration needs the same care, with additional attention to transaction history, unsupported features, deletes, generated columns, and downstream dependencies. A migration plan should inventory readers and writers before changing the table's ownership. The data migration planning guide can help organize dependencies, validation, and rollout decisions.

Observability around the table

Iceberg's snapshots tell you what table state was committed. They don't tell you whether a source delivered its expected records, whether business values are plausible, or whether a dashboard's key metric moved outside its normal behavior. Those checks belong in the surrounding data reliability system.

For example, a team can validate record rules after a snapshot commit, track the arrival time of each expected load, compare current metrics with historical behavior, and flag schema changes before BI models fail. The checks can run against the table in place, without copying the underlying data into a separate monitoring store.

Best Practices Performance Tuning and Troubleshooting for Production

Production Iceberg operations revolve around keeping metadata, files, snapshots, and policies healthy together. A table can remain transactionally correct while becoming expensive to query because it contains too many small files, stale snapshots, or overlapping layouts from several partition specifications.

Start with file management. Monitor small-file creation, compact compatible files, and choose write settings that produce practical file sizes for the engines serving the table. Compaction should follow workload evidence. A table used for frequent incremental reads may need a different maintenance rhythm from a table serving occasional analytical scans.

Metadata deserves its own monitoring. Track manifest growth, planning latency, snapshot accumulation, and failed commit patterns. Expire snapshots according to recovery and audit requirements, then clean orphan files only after confirming that no active snapshot or external process still depends on them.

Operational rule: Maintenance isn't housekeeping. It's part of the table's query and recovery design.

Troubleshooting becomes more systematic when symptoms map to layers:

  • Slow planning: Inspect manifest counts, metadata growth, and catalog response time before blaming the query engine.

  • Slow scans: Check pruning effectiveness, file sizes, data distribution, and whether evolved partition specs create uneven layouts.

  • Missing records: Compare the expected ingestion window with the committed snapshot and validate source delivery separately.

  • Schema failures: Identify the writer that changed the schema, then check reader compatibility and downstream assumptions.

  • Commit conflicts: Review concurrent writers, retry behavior, and maintenance jobs that may be competing for the same table.

  • Unexpected storage growth: Examine retained snapshots, delete files, failed writes, and orphaned data files.

Format upgrades require a rollout plan. Iceberg version changes are opt-in table by table, so older tables can coexist with newer ones. The 2025 Apache Iceberg ecosystem survey reported that 78.6% of respondents use Iceberg exclusively among open table formats, while an independent enterprise study reported 58% using Iceberg for business-critical analytics, 95% using or planning to use it for AI/ML, and 79% moving or planning to move remaining data to Iceberg within 12 months. These figures come from the 2025 State of the Apache Iceberg Ecosystem survey summary, and they signal momentum, not proof that every organization has solved mixed-version operations.

The next planning concern is Iceberg v3, with capabilities such as row lineage, deletion vectors, and new logical types. Test readers and writers together, define compatibility gates, and upgrade representative tables before broad rollout. The format may be open, but an estate can still accumulate hidden reliability debt if catalogs, engines, and governance policies evolve at different speeds.

The catalog should be selected with equal care. Iceberg provides ACID behavior, schema evolution, and time travel, while lineage, classification, access control, and unified audit trails depend on the catalog and governance stack. Databricks announced Managed Iceberg, Iceberg v3, and Foreign Iceberg as generally available in Unity Catalog on May 28, 2026, while Snowflake embedded Polaris into Horizon Catalog for Iceberg REST interoperability, as described in the Apache Iceberg specification materials. Those ecosystem moves reinforce the central decision: cross-engine openness depends not only on the table files, but also on who controls metadata access and policy enforcement.

digna runs inside your own environment and combines schema tracking, timeliness monitoring, in-database validation, anomaly detection, and data platform observability around critical tables and pipelines. Visit digna to see how your team can monitor Iceberg-based analytics without moving production data.

Most production Iceberg incidents turn out to be value problems wearing a metadata costume, so plan for data quality management alongside the migration.

Frequently asked questions

What problem does Iceberg solve?

It stops different engines disagreeing about what a table contains. Without a table format, one engine can read a directory before a batch is fully committed, another can interpret a partition differently, and a third can still see an older schema — while every file involved is perfectly valid.

How is an Iceberg table structured?

Iceberg keeps a chain of metadata: a catalog entry points to a metadata file describing the current snapshot, that snapshot points to a manifest list, and manifests enumerate data files with their statistics. Every commit writes new metadata rather than mutating the old, which is what makes rollback cheap.

Does Iceberg support concurrent writers?

Yes, through optimistic concurrency. Each writer prepares its changes and then attempts to swap the table's metadata pointer; if another commit landed first, the writer retries against the new state. Conflicting writes to the same files fail loudly instead of silently overwriting one another.

Which engines work with Iceberg?

Spark, Trino, Flink, Dremio, Snowflake and BigQuery among others read Iceberg, though feature coverage differs — some engines read but do not write, and newer capabilities land unevenly. Confirm the specific operations you need are supported by every engine in your stack, not just the one you test with.

What usually goes wrong with Iceberg in production?

Small files and unexpired snapshots, in that order. Streaming or frequent batch writes produce many small data files and a long metadata chain, which slows planning until compaction and snapshot expiry run on a schedule. Neither degrades correctness, so the symptom appears as creeping query latency.

✦ Generato con l'intelligenza artificiale

Condividete su X
Condividete su X
Condividete su Facebook
Condividete su Facebook
Condividete su LinkedIn
Condividete su LinkedIn

Il team dietro la piattaforma

Un team con sede a Vienna di esperti di AI, dati e software, supportato

da rigore accademico ed esperienza enterprise.

Il team dietro la piattaforma

Un team di esperti di IA, dati e software con sede a Vienna, forte di rigore accademico ed esperienza aziendale.

Prodotto

Integrazioni

Risorse

Azienda

INDEXED BYIndexerNow INDEXED BYIndexerNow