• novità

    • Release 2026.06 - Portiamo la data observability nel vostro codice

  • novità

    • Contribuite al futuro dell’innovazione in IA e dati

Open Table Format Guide: Iceberg, Delta, and Hudi Compared

|

8

min di lettura

The most popular advice about an open table format is to choose one that prevents vendor lock-in. That advice is incomplete. File portability matters, but the deeper change is that table meaning, transaction history, schema intent, and query-planning metadata move into a shared layer above objects in cloud storage.

That shift turns a directory of Parquet or ORC files into a table that several engines can discover, read, update, and evolve. It also creates a new responsibility: metadata must be governed. Iceberg, Delta Lake, and Hudi can make storage behave more like a database, but they don't decide whether a late snapshot violates an SLA, whether a new column breaks a downstream model, or whether a valid-looking dataset contains an abnormal business pattern.

This guide starts with the conceptual model, compares the three major formats, and then follows the metadata into production operations. The central question is simple: what control plane keeps an open table trustworthy after the commit succeeds?

Table of Contents

Why Open Table Formats Change the Lakehouse Game

An open table format isn't merely a common way to read Parquet without depending on one vendor. It places table semantics in metadata that can live alongside the data in object storage. Those semantics include atomic commits, versioned snapshots, schema evolution, partition information, and file-level statistics.

Without that layer, object storage mostly gives you files and folders. A query engine may infer a table from a directory, but directory structure doesn't provide a reliable answer to basic questions: which files belong to the current version, which write completed, what schema was active yesterday, or which files can safely be skipped for a filter.

Open table formats address those questions through shared metadata. Apache Hudi was created at Uber in 2016, Delta Lake was introduced by Databricks in 2017–2018 and open-sourced in 2019, and Apache Iceberg was born at Netflix in 2018 and donated to the Apache Software Foundation in 2019. These separate efforts converged on the same problem, making data lakes behave more like databases without moving data out of object storage (historical overview of open table formats).

From directories to first-class tables

The practical result is a table that can support batch reads, streaming writes, updates, and deletes on the same underlying storage. Engines don't need to treat every file as an independent source of truth. They can resolve a consistent snapshot, inspect metadata, and plan work against the relevant subset of files.

That changes engineering design in several ways:

  • Concurrent writes become manageable: A commit can succeed as a complete metadata change or fail without exposing a partial table state.

  • Multiple engines can cooperate: Spark, Trino, Flink, warehouses, and other readers can use the table's metadata contract instead of relying only on physical paths.

  • Pipelines become less engine-specific: Teams can separate storage semantics from the compute engine that performs a transformation.

  • Tables become catalog objects: Ownership, discovery, access, and lineage can be organized around a table rather than a folder.

A Delta table, for example, contains data files and a transaction log. Periodic checkpoints compact that log so engines can identify the active snapshot and relevant partitions without enumerating every object, even at very large scale (technical description of Delta Lake's transaction design).

Architectural rule: An open table format gives you a durable description of table state. It doesn't automatically give you durable governance of everything that state means to the business.

That distinction matters for lakehouse design. A useful lakehouse data quality model must sit above storage commits and connect structural changes to ownership, validation, freshness, and downstream impact.

Core Primitives Explained Without the Jargon

Think of object storage as a large postal mailroom. The floor holds boxes containing raw Parquet or ORC files. The boxes are durable, but the floor alone doesn't tell the clerk which boxes belong to today's delivery or whether a newly arrived box follows the agreed envelope format.

An open table format adds the clerk's records. Each record describes the table's current state and helps a query engine find the right boxes without opening every one.

A diagram illustrating the five core primitives of open table formats including data files, schema, snapshots, partitioning, and statistics.

The five records the clerk needs

Data files are the boxes on the mailroom floor. They contain the actual rows, usually in columnar formats such as Parquet or ORC. The role of Parquet in analytical storage is separate from table management. Parquet stores data efficiently, while the open table format explains how those files form one evolving table.

Schema is the rulebook for envelope shape. It defines column names, data types, and structural expectations. Schema evolution lets a table change over time without forcing every historical file to be rewritten immediately, but the permitted change still needs governance.

Manifests are detailed inventories. A manifest entry can identify a data-file path, its partition values, and column statistics. In Iceberg, metadata files define the table, manifest lists define snapshots, and manifest files list data files with statistics used during pruning (Iceberg architectural comparison).

Snapshots are point-in-time views of the clerk's records. A reader can ask for the table as it existed at a prior snapshot and receive a consistent answer, even while new files arrive. Iceberg uses immutable metadata files and snapshot history, with each commit creating a new metadata version (Iceberg specification).

ACID transactions are the clerk's commit rules. A metadata update either becomes visible as a complete operation or doesn't become the current table state. That protects readers from seeing half of a write and gives writers a way to coordinate changes.

Why partitioning and statistics affect performance

Partitioning is the shelf layout. If rows are grouped by a value that commonly appears in filters, an engine can skip entire shelves. Statistics add finer guidance, such as minimum and maximum values for a column inside a file, so the planner can avoid files whose ranges cannot match the query.

This is why performance depends on metadata planning rather than only on raw file speed. A well-organized table lets the engine reduce I/O before it scans the data. A poorly organized table can remain technically correct while forcing broad scans.

Schema evolution deserves the same practical caution. Adding a column may be compatible with existing readers, but dropping or changing a type can affect dashboards, models, and data contracts. The format records structural state. Your platform still needs to decide whether a proposed change is acceptable.

How Iceberg, Delta Lake, and Hudi Came to Exist

Iceberg, Delta Lake, and Hudi grew from separate teams confronting the same architectural gap: files on object storage aren't tables. Each project added metadata and commit behavior, while prioritizing different workloads and operating environments. Their history is therefore a history of metadata governance, not file layout.

Apache Hudi emerged at Uber in 2016, as teams needed lake storage for incremental processing and mutable data. Its design centers on upserts, deletes, record-level indexing, and a commit timeline. Those choices suit pipelines that repeatedly apply changes from operational systems.

Databricks introduced Delta Lake in 2017–2018 and open-sourced it in 2019. Delta stores transaction history in a log beside the data files. Sequential JSON or Parquet entries under _delta_log record operations, allowing engines to reconstruct table state and query prior versions. See this architecture comparison of Iceberg, Delta Lake, and Hudi for a concise view of their designs.

Netflix developed Iceberg in 2018 and donated it to the Apache Software Foundation in 2019. Iceberg focused on immutable metadata, snapshot isolation, hidden partitioning, and engine-neutral table access. Its metadata structure separates table definitions, snapshot references, and file-level manifests, supporting partition evolution and query pruning. Our Apache Iceberg open table format overview provides more context.

A timeline graphic showing the convergent evolution of open table formats including Apache Hudi, Delta Lake, and Iceberg.

These formats converged because they address the same foundation through different metadata models. The decision is how workloads write data, which engines must interoperate, how catalogs coordinate access, and which team will own table maintenance.

That history also clarifies the remaining operational gap. Hudi's record-oriented strengths, Delta's transaction-log approach, and Iceberg's manifest-based model provide useful table primitives. None independently defines enterprise-wide lineage, business definitions, access policy, anomaly response, or change approval. An in-database observability layer such as digna can connect those controls to schema changes, timeliness, and anomalous behavior.

Comparing Iceberg, Delta Lake, and Hudi Side by Side

Platform architects should compare formats by workload shape and engine mix, not by isolated feature claims. All three can provide transactional table behavior, schema handling, version history, and metadata-driven query planning. Their differences become clearer in how they represent state and optimize writes.

Dimension

Apache Iceberg

Delta Lake

Apache Hudi

Metadata model

Immutable metadata files, manifest lists, manifests, and snapshot history

Flat transaction log under _delta_log, with JSON or Parquet entries and checkpoints

Timeline-based commits with metadata and indexing structures oriented toward record changes

Partitioning style

Hidden partitioning and partition transforms, with partition evolution handled as a metadata change

Partition columns plus ecosystem-specific layout and optimization features

Partition columns, record-level indexes, and engine-supported organization for mutable workloads

Schema evolution

Explicit schema and partition metadata, with evolution designed to avoid tying queries to physical partition columns

Schema and transaction metadata managed through the Delta protocol and surrounding platform controls

Schema and commit timeline support changes for ingestion and update-heavy pipelines

Concurrency guarantees

Snapshot-based commits and isolation through immutable metadata versions

Transaction-log commits and optimistic coordination around table state

Timeline commits designed for inserts, updates, and deletes

Ecosystem fit

Strong option for heterogeneous engine environments and catalog-oriented architectures

Natural fit for Databricks and Spark-centered platforms, with broader interoperability depending on protocol support

Strong option for upsert-heavy ingestion, incremental processing, and workloads that benefit from record-level indexing

Apache Iceberg

Iceberg is often attractive when many engines must share tables and physical partition columns shouldn't leak into query logic. Hidden partitioning lets the format manage transforms while queries refer to logical columns. Its specification also supports partition evolution as a metadata-only change that doesn't eagerly rewrite existing data files (Iceberg documentation).

The trade-off is metadata management. Manifest files and snapshot structures provide rich planning information, but small or frequently changing tables can accumulate metadata that requires deliberate maintenance. Iceberg is a strong fit when interoperability, partition evolution, and snapshot-oriented governance outweigh the simplicity of a single-vendor stack.

Delta Lake

Delta's transaction log is easy to reason about for teams already operating Spark and Databricks. The log records table actions, and checkpoints make state reconstruction more efficient at scale. That integration can shorten the path from existing Spark pipelines to transactional lakehouse tables.

The candid limitation is coupling. Delta can work across a broader ecosystem, but the smoothest experience often depends on Databricks protocols, runtimes, catalogs, and optimization practices. A platform that expects independent engines to write and govern the same tables should verify those paths rather than assuming compatibility from the file layout alone.

Apache Hudi

Hudi is compelling for change-data-capture pipelines and mutable tables. Its record-level indexing approach can direct updates and deletes toward the relevant file groups, reducing the need to search broadly for the target records. That strength comes with more format-specific operational concepts, including indexing, compaction, and commit timelines.

Choose Hudi when upsert behavior is central to the workload and the team is prepared to operate its write and maintenance model. Choose Delta when Spark and Databricks are the center of gravity. Choose Iceberg when engine neutrality, catalog coordination, and evolving partition strategies are primary requirements.

Decision test: List the engines that must read and write the same table, then list the mutations, governance controls, and rollback behaviors those engines require. The format that fits that intersection is more important than a generic portability claim.

Where Open Table Formats Stop Solving the Problem

An open table format can tell you that a commit succeeded. It can't tell you whether the committed records satisfy a business rule.

That boundary is deliberate. The format layer manages table structure and state, not the meaning of every record or the service-level expectations around delivery. It won't automatically validate that a transaction has a permissible status, alert that a snapshot arrived late, identify an unexpected distribution change, or reconcile a schema change against every downstream contract.

The governance gap above the table

Schema evolution shows the problem clearly. A table format may allow an additive change without breaking existing queries, but a production team still needs to ask who approved it, which consumers depend on the affected column, whether a type widening changes model behavior, and whether an audit trail exists.

A 2026 Databricks explainer highlights that easy schema changes can encourage direct production edits without sufficient testing, while catalogs become important for authoritative snapshots and access control (Databricks discussion of open table formats). Independent 2026 coverage also describes a gap between table-level structure and catalog-wide lineage, business glossaries, classification, and cross-engine policy management (analysis of modern data architecture).

Teams often fill that gap with disconnected tools:

  • Catalog workflows: Asset discovery and ownership live in one system.

  • Validation jobs: Business checks run in notebooks or orchestration tasks.

  • Freshness monitors: Arrival alerts depend on schedules and pipeline metadata.

  • Incident handling: Analysts and engineers use separate dashboards to investigate impact.

That fragmentation makes a metadata incident harder to follow. A new column appears in an Iceberg manifest, a downstream model continues running, a dashboard changes, and no single view connects the structural event to the resulting business risk.

Governance can move upward, not disappear

Open formats reduce storage lock-in, but organizations can still create lock-in above the files through catalogs, access-control systems, lineage conventions, and operating practices. The challenge becomes sharper in finance, healthcare, telecom, and public-sector environments, where teams need consistent controls across engines and evidence for important changes.

A control plane must therefore observe more than file state. It should connect schema, timeliness, record validity, anomalies, ownership, and downstream dependencies while keeping the data in the environment where the organization already governs it.

A data lake monitoring approach should begin at the table boundary, then follow the metadata and data behavior outward into pipelines, consumers, and business controls.

How digna Closes the Operational Gap

digna can sit above Iceberg, Delta Lake, and Hudi as an in-database observability layer. The useful mental model is not another table format. It's a set of operational checks that turns table metadata and snapshot behavior into signals for engineers, data owners, and governance teams.

Start with structural change

Schema Tracker watches structural metadata and detects column additions, removed columns, and data-type modifications. For an open table, that means relating changes in Iceberg metadata files, Delta transaction-log entries, or Hudi commit information to the owners and consumers that depend on the table.

A schema event shouldn't automatically become an incident. The important question is whether the change is compatible with the table's contracts. A new nullable attribute may be acceptable, while a removed identifier or incompatible type change may require approval and downstream testing.

Add delivery and record behavior

Timeliness compares actual snapshot arrival with expected delivery patterns and service-level expectations. It can flag late-arriving partitions, missing loads, and early deliveries, which helps teams distinguish a successful commit from a successful data product delivery.

Data Anomalies profiles new snapshots to surface unusual volume changes, null behavior, and distribution shifts. That complements file-level statistics. Metadata can help a query engine prune files, but observability asks whether the newly committed data looks normal for its business context.

Data Validation applies record-level rules and schema contracts. Those checks can enforce conditions such as valid relationships, permissible values, or audit requirements before consumers treat a snapshot as ready.

A diagram illustrating the Digna platform features for managing open table formats like Iceberg, Delta, and Hudi.

Connect incidents to platform health

Data Platform Observability brings pipeline health, lineage, freshness, consumption, and platform behavior into one operational view. If an upstream Iceberg table changes schema, the team can trace that event to affected dashboards or models instead of searching separately through storage logs, orchestration history, and BI alerts.

digna's in-database execution model keeps metric computation and analysis inside the customer's own databases. Its deployment model supports private cloud and on-premises environments, which can matter when regulated data can't be copied into an external monitoring service.

The architectural boundary remains clear. Iceberg, Delta, or Hudi owns table state and metadata. digna monitors whether that state is structurally safe, timely, valid, and behaving as expected.

A Practical Path to Adoption and Migration

Migration works best as a controlled sequence, not a bulk conversion. Treat each table as a product with a storage format, a catalog identity, a set of consumers, and operational expectations.

Phase one, inventory the estate

List current tables, file formats, write patterns, partition schemes, engine dependencies, and downstream consumers. Identify which tables receive append-only data, which require updates or deletes, and which are read by more than one engine.

This inventory exposes the actual decision criteria. A table used by Spark alone has a different migration path from one written by Flink, queried by Trino, and consumed by a warehouse.

Phase two, choose the target

Select Iceberg, Delta Lake, or Hudi based on engine fit, mutation patterns, partitioning needs, catalog behavior, and the team's operating skills. Finance teams may prioritize transactional consistency and schema enforcement. Retail pipelines may favor hidden partitioning and upserts from streaming CDC. SaaS analytics platforms may place greater weight on multi-engine reads.

The format is only one choice. Decide which catalog will define table identity, how engines will discover tables, who approves schema changes, and how rollback works.

Phase three, wire controls before cutover

Connect the catalog, validation workflows, ownership records, and observability before production traffic moves. Run side-by-side reads from the legacy table and the Iceberg, Delta, or Hudi target. Compare schema and row-level outcomes with Data Validation, then set gates for freshness and anomaly behavior.

A data migration planning framework should include rollback criteria, shadow-read duration, consumer sign-off, and evidence retention. Don't wait for the first failed dashboard to discover that nobody owns the new table.

Phase four, migrate table by table

Convert a limited table, validate reads and writes, monitor metadata growth, and observe query behavior. Keep the legacy path available until parity and operational thresholds are met, then move consumers in stages.

The format choice matters, but metadata discipline matters more. A well-governed Hudi table can outperform a poorly maintained Iceberg deployment for its workload, just as a carefully operated Delta estate can be a poor fit for an organization that requires independent engines and catalogs.

A checklist infographic outlining the four phases of data lake adoption and migration, from inventory to optimization.

Choose the format that matches your workload, but design the governance layer at the same time. That is what turns open storage into a dependable lakehouse table.

digna monitors schema changes, snapshot timeliness, record validity, and anomalous data behavior inside your own environment, giving teams an operational layer above Iceberg, Delta Lake, and Hudi. Visit digna to evaluate how its modular observability and validation capabilities can support your open table migration.

Where the format stops, business-level checks begin — business monitoring watches whether the numbers a table produces still behave.

Frequently asked questions

Is avoiding vendor lock-in the main reason to use an open table format?

That reason is incomplete. File portability matters, but the bigger shift is that table meaning — transaction history, schema intent and query-planning metadata — moves out of a single engine and into files anyone can read. Lock-in is the symptom; where the table's definition lives is the actual change.

What primitives do Iceberg, Delta Lake and Hudi share?

All three keep an immutable log of commits, a manifest describing which files belong to each version, per-file statistics for pruning, and a schema record that evolves independently of the data files. The differences lie in how each one writes and compacts that structure, not in the underlying concept.

How do the three formats differ in practice?

Iceberg has the broadest engine support and the most engine-neutral design. Delta Lake runs deepest inside Spark and Databricks estates. Hudi optimises for frequent upserts and incremental consumption. The honest comparison is about which write pattern each was built around, not which lists more features.

Where do open table formats stop solving the problem?

At the boundary between structure and meaning. The format can guarantee every engine reads the same committed rows under the same schema; it cannot tell you those rows are complete, arrived on time, or hold sane values. That gap is where quality monitoring belongs.

What does a realistic adoption path look like?

Pick one table with a clear owner and a known query pattern, convert it while the original keeps running, and compare results over a full business cycle before switching consumers. Assign maintenance — compaction, snapshot expiry, catalog upgrades — before the pilot ends, or it becomes nobody's job.

✦ Generato con l'intelligenza artificiale

Condividete su X
Condividete su X
Condividete su Facebook
Condividete su Facebook
Condividete su LinkedIn
Condividete su LinkedIn

Il team dietro la piattaforma

Un team con sede a Vienna di esperti di AI, dati e software, supportato

da rigore accademico ed esperienza enterprise.

Il team dietro la piattaforma

Un team di esperti di IA, dati e software con sede a Vienna, forte di rigore accademico ed esperienza aziendale.

Prodotto

Integrazioni

Risorse

Azienda

INDEXED BYIndexerNow INDEXED BYIndexerNow