• new

    The major Release 2026 is live - Bringing Data Observability Into Your Code

  • new

    Contribute to the Future of AI & Data Innovation

  • new

    • Release 2026.06 - Bringing Data Observability Into Your Code

  • new

    • Contribute to the Future of AI & Data Innovation

What Is Open Table Format

|

10

min read

You're probably here because your lake already looks like a table in one tool, but behaves like a pile of files in another.

A finance dashboard changes after the fact. A backfill fixes one query and breaks another. Spark writes succeed, Trino reads stale data, and nobody can answer a basic question: which version of this dataset is the correct one? That's the point where people stop asking “which file format should we use?” and start asking what is an open table format, really?

The short answer is simple. An open table format is the metadata layer that makes object storage behave like a database table. The longer answer matters more, because the format itself is only half the story. The other half is how your team operates it: compaction, schema control, catalogs, and observability.

Table of Contents

The Moment a Data Lake Stops Being Trustworthy

A common failure starts.

Finance checks yesterday's revenue report and sees that it no longer matches the dashboard. No pipeline is red. No orchestrator shows a failed task. The files are still in object storage. But somewhere between a streaming append and a maintenance job, readers picked up a partial view of the table.

Why raw storage breaks in subtle ways

Object storage is very good at holding files. It is not, by itself, a table system.

If you store Parquet files in a folder and call that folder revenue, you still haven't answered the questions analysts care about:

  • Which files belong to the table right now

  • Which write produced the current version

  • Which schema should a reader use

  • What did the table look like last week

  • Which files are safe to ignore during a query

Without that metadata layer, each engine has to infer state from directory listings and file names. That works for a while. Then someone rewrites partitions, compacts files, or changes a column type, and readers start seeing different truths from the same storage path.

Raw files can store data correctly and still produce unreliable analytics.

The missing layer

An open table format adds the missing record of table state alongside the data itself. Instead of treating storage as “whatever files happen to be under this prefix,” it records a governed view of the table: current files, historical snapshots, schema, and layout metadata.

That's why teams that invest in data lake monitoring often discover the same pattern. The issue usually isn't just a broken file write. It's that the platform lacks a reliable way to describe table state across readers and writers.

The difference feels small until your first silent inconsistency. After that, it becomes foundational.

What an Open Table Format Actually Does

A new pipeline lands replacement files for revenue. The Spark job finishes. An analyst refreshes a dashboard in Trino. A machine learning feature job reads the same table through another engine. All three queries hit the same storage path, but they do not always see the same answer.

That gap is what an open table format is built to close.

A diagram illustrating how an open table format organizes data files, metadata, and snapshots in storage.

It sits above the files and defines table state

Open table formats are often confused with file formats because both show up in the same storage bucket. They solve different problems.

Parquet, ORC, and Avro define how a single file stores rows and columns. An open table format defines how a collection of files becomes one logical table that many engines can read and write safely. Databricks describes open table formats as a metadata layer on top of Parquet, ORC, or Avro in object storage that adds table features such as ACID transactions, schema evolution, and time travel in its overview of open table formats.

If that distinction still feels blurry, this guide to what Parquet is and how it stores columnar data helps. Parquet answers "how is this file encoded?" A table format answers "which files make up the table, and which version should readers trust?"

The format acts like the table's control plane

A useful mental model is a Git history for data files, with stricter rules around concurrency and schema. The table format keeps a record of the current state and the path that produced it. Query engines do not have to scan folders and guess. They can read the table metadata and get an exact answer.

In practice, that metadata layer handles four jobs.

  1. Declare file membership
    The format records which files belong to the table right now. If compaction rewrites ten small files into one larger file, readers follow the metadata pointer to the valid set instead of mixing old and new data.

  2. Publish snapshots atomically
    A write becomes visible as a new committed snapshot. Readers see the old version or the new version, not an in-between state where half the files have changed.

  3. Track schema as table metadata
    Column names, types, field IDs, and allowed changes live in the table definition, not only in whatever a writer happened to emit. That matters once different engines start renaming columns, widening types, or adding nested fields.

  4. Store layout and pruning metadata
    Partition values, file statistics, and sometimes delete metadata help engines skip irrelevant files and plan queries efficiently.

The operational layer is the part teams usually underestimate

The feature list is the easy part. The operating model is where open table formats earn their keep.

Once a format owns table state, your team has to decide who owns compaction, who approves schema changes, and which catalog is the source of truth. Those are production questions, not product brochure questions. If nobody owns compaction, small files pile up and query performance drifts. If every writer can change schema freely, one bad deployment can break downstream readers. If Spark uses one catalog view and Trino uses another, you are back to multiple truths with better branding.

The format gives you the mechanisms. Your platform team still needs rules.

Why the change matters in daily work

Without a table format, writing data often means "files were uploaded." With a table format, writing data means "a new table version was committed and published."

That is a stronger contract for every downstream system. Batch jobs, BI queries, and streaming readers can agree on table state. Engines can interoperate through shared metadata instead of storage path conventions. Data engineers can reason about maintenance work, such as compaction or partition evolution, without treating every rewrite as a potential outage.

An open table format, then, is the layer that turns object storage into a table system people can operate with confidence. The file format still matters. The hidden win is that the metadata gives your team a controlled way to keep table state consistent as writes, schemas, and engines multiply.

How the Format Evolved From Hive to Lakehouse

The history matters because each format was created to fix a specific pain, not to win a feature matrix.

A useful timeline from the history and evolution of open table formats places Apache Hive in 2008, Apache Hudi in 2016, Apache Iceberg in 2017, and Delta Lake in 2017 to 2019. That sequence shows how quickly the lakehouse table layer matured once teams hit the limits of early data lake patterns.

A timeline chart illustrating the evolution of data storage formats from Apache Hive to modern open lakehouse solutions.

Hive solved one era's problem

Hive gave early Hadoop teams a way to treat files on HDFS as tables. Its partition model leaned heavily on directory structure. That fit the assumptions of the time.

Those assumptions aged badly in cloud object storage. Directory-style partitioning became brittle. Listing and rename behavior didn't feel like HDFS. Long-lived tables became awkward when partition strategies changed.

The next wave fixed concrete failure modes

Each newer format responded to a different operational weakness.

  • Apache Hudi emerged in 2016 to support workloads that needed record-level upserts and incremental processing in data lake environments.

  • Apache Iceberg arrived in 2017 to address problems around large analytic tables, metadata scaling, and partition evolution.

  • Delta Lake took shape in 2017 to 2019 as a transaction-log-based way to bring reliable ACID-style behavior to open storage.

By that point, the industry had moved beyond “how do we put files in a lake?” to “how do multiple engines share trustworthy tables on cheap storage?”

The lakehouse name came after the engineering shift

The architectural shift came first. The lakehouse label followed.

If you want the broader stack context around formats, catalogs, engines, and governance, this guide on what a lakehouse is and how to maintain data quality is the bigger-picture view. The important point here is narrower: open table formats became the layer that made lakehouse claims technically credible.

By 2025 to 2026, that layer had matured further. The same timeline notes that Iceberg's v3 table spec was ratified in 2025, and Iceberg 1.10 and 1.11 brought stable support for features including deletion vectors, variant types, geospatial types, and nanosecond timestamps in that later phase of the format's evolution.

Iceberg, Delta Lake, and Hudi Compared

A format decision usually looks simple during evaluation. Then production starts, a streaming job falls behind, small files pile up, two engines disagree on schema changes, and the question appears: who owns table maintenance, and how predictable is that work?

Apache Iceberg, Delta Lake, and Apache Hudi are all production-grade table formats. The useful comparison is not just feature parity. It is how each format organizes metadata, coordinates writes, and assigns responsibility for operational chores such as compaction, cleanup, and cross-engine compatibility.

The biggest difference is metadata architecture

The storage files may all sit in object storage, but each format keeps the table's "table of contents" differently. Dremio's overview of Iceberg, Delta Lake, and Hudi architecture is a good reference point:

  • Delta Lake records table changes in a sequential transaction log under _delta_log.

  • Iceberg tracks table state through snapshot metadata and manifest files that describe which data files belong to each version.

  • Hudi keeps a timeline of actions, including commits and maintenance operations such as compaction, clustering, and cleaning.

That design choice affects more than read performance. It changes how quickly metadata grows, how safely multiple engines can write, how partition changes are handled, and how much routine care the table needs from your team.

Iceberg vs Delta Lake vs Hudi at a Glance

Dimension

Apache Iceberg

Delta Lake

Apache Hudi

Metadata design

Snapshot metadata with manifest lists and manifests

Sequential transaction log in _delta_log, with checkpoints to limit replay cost

Timeline of immutable instants covering writes and table services

Partitioning model

Hidden partitioning and partition evolution, so the logical partition scheme can change without rewriting older files

Partition behavior is managed through log-tracked table state and writer conventions

Designed around ingestion-heavy layouts, with services to reorganize files over time

Concurrency style

Snapshot isolation model centered on table versions and metadata swaps

Log-based commit protocol with optimistic concurrency patterns

Timeline-based commit protocol, often paired with background services

Best fit

Shared analytic tables across multiple engines

Spark-centered platforms, especially where Delta-native tooling is part of daily operations

Frequent upserts, change capture, and incremental processing

Main trade-off

Strong interoperability, but catalog discipline matters and cross-engine write rules need to be explicit

Operational patterns are often easiest where Spark and Delta-aware tooling set the standard

Write flexibility is strong, but compaction and clustering ownership need active planning

Interoperability posture

Broad support across engines and catalogs. The Apache Iceberg specification is one reason many teams treat it as the most neutral shared-table option

Strong Spark heritage, with improving compatibility across the wider ecosystem

Interoperability has improved, but operating assumptions still tend to be more ingestion-focused than engine-neutral

A better decision lens: who will operate the table

If a team compares only checkboxes, all three can look close. In practice, they create different on-call work.

Iceberg often fits organizations that want one table to be read by several engines and governed through a catalog layer. The benefit is portability. The cost is discipline. Schema changes, sort order choices, file sizing, snapshot retention, and catalog consistency all need clear ownership, especially once more than one compute engine can write.

Delta Lake often fits platforms that are already centered on Spark workflows and Delta-aware tooling. The operational path can be straightforward because the transaction model and surrounding tooling are tightly aligned. The trade-off is that interoperability strategy matters more over time if the platform expands beyond its original engine assumptions.

Hudi often fits ingestion-heavy systems where record-level updates and incremental consumption are part of the core design. That strength comes with more visible table services. Compaction, clustering, and cleanup are not side details. Someone has to schedule them, monitor them, and decide when write latency matters more than read layout.

This is also where reliability work stops being separate from quality work. A late compaction job, an unreviewed schema change, or an engine writing with the wrong catalog settings can create incidents that analysts experience as "bad data." Teams running Spark-first platforms often evaluate format choice alongside Databricks data quality management practices, because table correctness and data quality checks meet in the same production workflow.

The wrong format rarely fails during a demo. It fails during the maintenance task your team assumed would take care of itself.

ACID, Time Travel, and Schema Evolution Explained

A table format starts to feel real the first time something goes wrong in production. A pipeline writes half its files, fails on the rest, and an analyst refreshes a dashboard at exactly the wrong moment. Without transaction control, the lake exposes that mess directly. With an open table format, readers keep seeing the last valid table state until the new one is fully committed.

ACID means the table has a clear commit point

The practical benefit of ACID in a lakehouse is simple. Readers do not see half-finished work.

A diagram explaining ACID transactions, time travel, and schema evolution within open table format architecture.

Under the hood, these formats publish changes through metadata commits. Data files may already exist in object storage, but they do not become part of the table until the new metadata version is accepted. That gives you behavior data teams expect from a database, even though the underlying storage is still files in a lake.

In practice, that means:

  • a batch job can publish a full refresh as one atomic commit

  • a streaming job can keep appending while readers stay on a consistent snapshot

  • a failed write can leave orphaned files behind, but not become table truth

  • concurrent writers need conflict rules, retry logic, and catalog coordination to avoid stepping on each other

That last point is easy to underestimate. ACID is not only a reader feature. It is also a control mechanism for multiple writers, and the operational details differ by format, engine, and catalog setup.

Time travel is how teams debug and reproduce data states

Time travel means the table can expose older committed versions, not just the current one. That sounds like a convenience feature until a report changes after a backfill, or a deletion job removes more records than intended.

If you have worked with historical data in Snowflake, the mental model is similar. You query a prior table state instead of restoring raw files from backup.

The useful question is not "does the format support time travel?" The useful question is "how long do we keep history, who controls retention, and what happens when storage cleanup runs?" Teams often celebrate historical reads during evaluation, then discover later that aggressive snapshot expiration removed the exact version they needed for an audit or incident review.

Time travel is only as reliable as the retention policy behind it.

Schema evolution separates controlled change from accidental breakage

Schema evolution lets a table change shape over time without forcing a full rewrite of every old file. That usually includes adding nullable columns, handling some type promotions, and in certain formats, supporting column renames through stable field identity instead of raw column position.

The feature sounds permissive. Production use should be the opposite.

A safe schema change process answers a few specific questions:

  • Who can approve a schema change in production

  • Which engines are allowed to write that new schema

  • Whether downstream readers can handle the change

  • How partition changes affect existing jobs and assumptions

  • Whether the catalog will interpret the new schema the same way across tools

That is where teams either gain flexibility or create long-running cleanup work. A column rename may be valid at the format level and still break BI extracts, dbt models, or ingestion code that referenced the old name. A type widening may pass write validation and still change join behavior or null handling downstream.

Schema evolution is a format capability. Schema governance is an operating decision.

The same pattern applies across ACID and time travel too. The format provides the mechanism. Reliability comes from ownership: who approves changes, who tunes retention, who cleans up files, and who decides which engines are trusted to write.

The Real Question After Adoption

A lakehouse usually feels healthy right after the first successful write. Queries run. Time travel works. Everyone checks the feature list and moves on.

Then the operating questions start. Which catalog is the source of truth. Which engines are allowed to write. Who runs compaction. Who approves schema changes that are technically valid but risky for downstream jobs. Those decisions shape day-to-day reliability more than the file format choice alone.

The catalog sets the rules of the road

An open table format defines how table metadata and data files fit together. The catalog decides how systems find that table, coordinate changes, and apply policy. In practice, the catalog works like air traffic control for the lakehouse. The planes may all be built correctly, but you still need one place coordinating takeoffs, landings, and route changes.

Uplatz's review of open table format convergence and divergence points to the growing role of translation and interoperability projects such as Apache XTable, Hudi native Iceberg output, and Delta Lake UniForm. It also notes a production reality that gets missed in format comparisons. Many enterprises run more than one format, and interoperability still depends heavily on engine support, catalog behavior, and cloud-specific constraints.

A five-point checklist titled The Real Question After Adoption emphasizing catalog selection over file formats.

That is why the same Iceberg, Delta Lake, or Hudi table can feel predictable in one setup and fragile in another. A REST catalog, Glue, Hive Metastore, Nessie, Polaris, or a vendor-managed control plane will differ on discovery, permissions, branching models, write coordination, and failure handling.

Compaction ownership becomes an operating model

Compaction sounds like background maintenance. In production, it is closer to garbage collection mixed with index tuning. Done well, it keeps query planning sane and read performance stable. Done poorly, it burns cluster time, fights with ingestion workloads, and creates cleanup work after failed rewrites.

Someone has to own:

  • Compaction cadence

  • Sort and clustering choices

  • Snapshot cleanup and retention

  • Rollback steps when maintenance makes things worse

  • Orphan file detection

Those are not format checkboxes. They are recurring jobs with failure modes, cost trade-offs, and clear ownership needs. A candid review at Shirokoff on open table formats makes that point directly by calling out metadata overhead, maintenance work, learning curve costs, and uneven governance across tools.

Reliability depends on who controls change

Production teams also learn that governance rarely lives in the table files alone. Column masking, row-level controls, audit trails, and PII tagging usually depend on the catalog and the query engines enforcing them consistently.

The teams that keep these tables dependable watch the metadata layer like an application control plane. They track snapshot age, commit failures, file counts, orphaned files, maintenance lag, and which writer changed table state last. Teams that only watch object storage health often miss the earlier warning signs. The bucket looks fine while the table becomes harder to trust.

A Practical Rollout Plan for Data Engineering Teams

A clean rollout starts small, but it shouldn't start vaguely. You want a pilot that exposes operational edges early.

Phase 1 picks the control plane

Choose the catalog first. The file layout matters, but the catalog decides how engines discover tables, coordinate writes, and enforce policies.

Evaluate options such as REST-oriented catalogs, Hive Metastore-style setups, and cloud-native catalogs by asking practical questions:

  • Which engines must read and write natively

  • Where will table ownership and permissions live

  • How will schema changes be reviewed

  • What happens if the catalog is unavailable

If you skip this step, you'll often rebuild it later under pressure.

Phase 2 sets the security and governance baseline

Before broad adoption, define the minimum operating contract.

That usually includes SSO, RBAC, encryption in transit and at rest, controlled service identities, and a policy for row or column restrictions where needed. It also includes lineage expectations, PII tagging, and approval paths for schema changes.

One useful tool option in this stage is digna, which runs inside the customer's own environment and monitors data behavior, timeliness, record validation, schema changes, and platform metrics across lakes, warehouses, and pipelines. For open table formats, that matters because reliability problems often show up first as stale snapshots, schema drift, delayed writes, or business metric anomalies rather than hard pipeline failures.

Phase 3 proves the write path under stress

Don't validate the architecture with a single daily batch.

Use a pilot dataset with multiple daily writers, at least one maintenance workflow, and one query engine outside the primary writer engine. Include a hidden-partition or layout-optimization job in the trial so your team has to handle snapshot changes, compaction effects, and rollback decisions under realistic conditions.

A good pilot should create a few uncomfortable questions. Who approves schema changes? Who gets paged for failed commits? Which readers are allowed during maintenance windows? Those are the right problems to surface early.

Phase 4 instruments the metadata layer

Storage monitoring alone won't tell you enough.

Track signals such as:

  • Snapshot age for freshness of published table state

  • Compaction lag so file sprawl doesn't degrade performance

  • Writer failure patterns across engines

  • Schema drift events before downstream dashboards break

  • Orphan file growth after retries or aborted operations

If you can't see the health of the table metadata path, you're operating mostly on hope.

Open table formats solve the table-layer problem, but they don't remove the need for disciplined operations. digna helps teams monitor schema changes, timeliness, anomalies, and platform behavior inside their own environment, which fits the failure modes that show up after lakehouse adoption. If your tables already span multiple engines and maintenance jobs, it's worth looking at how digna can help you watch the metadata-driven reliability layer, not just the files underneath it.

Teams rolling out on Databricks usually add observability for the Databricks platform in the same phase, so table-level guarantees and value-level checks land together.

Frequently asked questions

How did table formats evolve from Hive?

Hive defined a table as a directory layout, so partitions were folders and the file listing was the source of truth. That broke down on object storage, where listings are slow and eventually consistent. Modern formats moved the table definition into metadata files, removing the dependency on directory structure entirely.

What does ACID mean for a data lake table?

It means a write either becomes visible in full or not at all, and concurrent readers are never exposed to an in-progress change. On a lake that guarantee comes from swapping a metadata pointer atomically, not from a database transaction log, which is why it works over plain object storage.

Can you change a table's schema without rewriting it?

Yes. Adding, dropping, renaming or reordering columns is recorded in metadata and applied at read time, so existing data files stay untouched. The catch is that readers must implement the same rules — an engine with partial support can return nulls for a renamed column rather than raising an error.

Why does a lake behave like a table in one tool and a pile of files in another?

Because the guarantee lives in the metadata layer, and only engines that read that layer get it. A tool pointed straight at the underlying Parquet files bypasses the commit history entirely and sees whatever happens to be on storage, including files from a write that never committed.

What does a sensible rollout look like?

Convert one well-understood table first, keep the original writing in parallel, and move consumers gradually while comparing outputs. Agree who owns snapshot expiry and compaction before go-live, because unowned maintenance is the most common reason a successful pilot degrades over the following quarter.

✦ Generated with Artifical Intelligence

Share on X
Share on X
Share on Facebook
Share on Facebook
Share on LinkedIn
Share on LinkedIn

Meet the Team Behind the Platform

A Vienna-based team of AI, data, and software experts backed

by academic rigor and enterprise experience.

Meet the Team Behind the Platform

A Vienna-based team of AI, data, and software experts backed by academic rigor and enterprise experience.

Product

Integrations

Resources

Company

INDEXED BYIndexerNow INDEXED BYIndexerNow