• novità

    • Release 2026.06 - Portiamo la data observability nel vostro codice

  • novità

    • Contribuite al futuro dell’innovazione in IA e dati

Parquet File Format Explained: A Practical 2026 Guide

|

10

min di lettura

You're probably dealing with Parquet already, even if you didn't choose it yourself. A warehouse export lands in S3, your Spark job reads a lake table backed by Parquet files, or Athena keeps scanning datasets that someone upstream partitioned three different ways over the last year. Most days, it feels fast enough and invisible enough that nobody asks questions.

Then something breaks. A query that used to fly starts dragging. A new writer rolls out a feature one engine can read and another can't. A single renamed field turns into a week of downstream nulls. That's when the Parquet file format stops being a file extension and starts becoming an operational concern.

Table of Contents

Why Parquet Became the Default Columnar Format

A common analytics pattern looks like this: a team stores a wide sales table with many attributes, but most dashboard queries touch only a handful of columns. In a row-oriented export like CSV, the engine still has to chew through every row as text. In Parquet, the engine can focus on the columns the query needs. That difference is why people keep choosing it for analytical storage.

Parquet didn't appear by accident. It began as an open-source collaboration between Twitter and Cloudera, with its first release on 13 March 2013, Parquet 1.0 in July 2013, and by 27 April 2015 it had become a top-level Apache Software Foundation project, according to the Apache Parquet file format documentation. Apache's own description is still the cleanest one: Parquet is an open source, column-oriented data file format for efficient storage and retrieval.

Why columnar storage changed the default choice

In plain terms, Parquet stores values by column instead of by row. All the order_total values sit together. All the country values sit together. That gives query engines two advantages:

  • Selective reading: They can skip columns your SQL never references.

  • Better packing: Similar values tend to compress well because the format can apply encodings before compression.

That combination fits modern analytics engines and lakehouse table formats well. If you want a shorter primer before going deeper, digna has a useful Parquet overview.

Practical rule: If your workload is mostly scans, aggregations, and filters across large datasets, Parquet usually fits the access pattern better than text formats.

Why it spread so widely

Parquet became the connective tissue across Spark pipelines, warehouse exports, object-storage data lakes, and table formats like Iceberg, Delta Lake, and Hudi. Teams value it because it's open, typed, and broadly supported. They also like that it separates physical storage from query engine choice. A file written in one part of the stack can often be read elsewhere.

That “often” matters. Portability is real, but it isn't automatic once newer features and mixed engine versions enter production.

Engine

Native Parquet Support

Common Pattern

Spark

Yes

Batch transformation and lakehouse tables

BigQuery

Yes

External tables and file exchange

Redshift Spectrum

Yes

Querying data in object storage

DuckDB

Yes

Local analytics and ad hoc querying

Athena

Yes

Serverless scans over partitioned lake data

Inside a Parquet File Row Groups, Column Chunks, and Pages

A Parquet file makes more sense if you picture a common warehouse query. An analyst asks for three columns out of fifty, filtered to last week's data. Whether that query feels fast or sluggish depends heavily on how Parquet lays bytes out inside the file, not just on the SQL engine reading it.

A hierarchical diagram illustrating the structure of a Parquet file format, including row groups, column chunks, and pages.

Start with the row group

A Parquet file is divided into row groups. Each row group is a horizontal slice of the table, so it contains the same set of columns for a subset of rows.

Row group size shapes read behavior more than many teams expect. Earlier Apache Parquet guidance recommends large row groups, because larger groups usually create larger column chunks and favor sequential I/O. In production, that can help scan-heavy workloads. It also creates a trade-off. Very large row groups reduce metadata overhead, but they can make selective reads less nimble because the engine still plans work at the row-group level.

Inside each row group, Parquet stores exactly one column chunk per column in the schema, and each chunk is contiguous in the file, as defined in the parquet-format repository. That layout is a big part of why column pruning works. If a query needs order_total and order_date, the engine can ignore the bytes for customer_notes, device_model, and everything else.

Then zoom into the column chunk and page

A column chunk holds one column's values for one row group. That chunk is then split into pages, which are the smaller units Parquet encodes and compresses.

Pages matter because they are where storage decisions become concrete. Page metadata records which encoding and compression were used. If a dictionary page exists, it must appear first in the column chunk, and there can be at most one dictionary page per column chunk, according to the Apache Parquet column chunk documentation.

A practical mental model is:

  1. File contains one physical Parquet object.

  2. Row groups divide the rows into large blocks.

  3. Column chunks store one column for one row group.

  4. Pages store encoded and compressed slices of that chunk.

A spreadsheet analogy helps here. Row groups are like bands of rows across the sheet. Column chunks are the vertical strips for each field inside one band. Pages are the smaller packets inside each strip that the reader can decode piece by piece.

Why the footer determines read planning

The footer is where Parquet keeps the map. It stores the schema plus metadata that tells the reader where row groups and column chunks live. Before a query engine can make smart decisions about what to read, it usually needs that footer.

That design is excellent for analytical planning. It is weaker for point access.

If your workload is “scan a month of events and aggregate,” footer-first planning helps the engine skip work. If your workload is “fetch one record by ID,” the reader still has to find and parse the footer before it can locate the right region of the file. Apache Arrow's Parquet documentation also notes that Parquet data is decoded in chunks rather than operated on directly at the byte level, which is one reason single-row retrieval is a poor fit for the format.

This is one of the production gaps that basic Parquet explainers skip. Teams often hear “columnar equals fast” and then apply Parquet to every access pattern. That usually holds for batch analytics. It often breaks for request-driven lookups, CDC validation probes, or operational debugging flows where an engineer wants one row right now.

The footer also plays into interoperability problems later. Different engines can agree on the file existing and still disagree on newer metadata features, page indexes, or optional statistics. When that happens, the symptom is rarely dramatic corruption. It is usually quieter. Filters stop pruning as well as expected, one engine reads more bytes than another, or query latency drifts upward months after a writer upgrade.

For analytics, footer-first planning is a strength. For point lookups, it is overhead.

That is why Parquet works best when you treat it as an analytical file format, watch how your engines use its metadata, and investigate when “bytes scanned” or row-group pruning rates suddenly get worse.

Encoding and Compression Options That Actually Matter

Teams often blur encoding and compression, but they solve different problems. Encoding restructures values inside a page. Compression then shrinks the encoded byte stream. In Parquet, those layers stack.

Encodings change the shape of the data first

Parquet supports page-level encodings including PLAIN, RLE, DELTA_BINARY_PACKED, RLE_DICTIONARY, and BYTE_STREAM_SPLIT, and it supports compression codecs such as SNAPPY, GZIP, BROTLI, ZSTD, and LZ4_RAW, as summarized in the Parquet format reference.

What matters in production is the data shape:

  • Dictionary-style paths help when values repeat, such as status fields, country codes, or product categories.

  • RLE works well when repeated values appear in runs, especially after sorting.

  • Delta-style encodings are useful when values change incrementally, such as increasing IDs or timestamps.

  • Byte stream split can help some numeric workloads, but compatibility may lag across tools.

An empirical evaluation of columnar formats found that Parquet's aggressive dictionary encoding often delivered strong compression with reasonable decode performance across many data types, and its bit-packing plus RLE path typically decoded faster than more complex multi-algorithm approaches, according to the evaluation paper on columnar formats.

Compression should match your operational goal

A useful way to choose codecs is to decide what you're optimizing for:

  • Fast reads and broad compatibility: use Snappy.

  • Tighter storage with a balanced tradeoff: use ZSTD where your engine stack supports it comfortably.

  • Maximum squeeze with slower reads and writes: use GZIP for cold or less interactive datasets.

Short version: choose encoding for the column's value pattern, then choose compression for the storage-versus-CPU tradeoff.

Data Shape

Encoding

Compression

Why It Works

Repeated strings in logs

Dictionary or RLE_DICTIONARY

SNAPPY

Repeated values compact well and Snappy keeps decode overhead low

Sorted status flags or region codes

RLE

ZSTD

Long runs compress efficiently and ZSTD usually balances ratio with read cost

Incrementing IDs or ordered timestamps

DELTA_BINARY_PACKED

GZIP or ZSTD

Small step changes encode compactly before the codec runs

Mixed low-cardinality dimension columns

Dictionary

SNAPPY or ZSTD

Similar values benefit from dictionary lookup and remain portable

Floating-point heavy analytical features

BYTE_STREAM_SPLIT where supported

ZSTD

Can improve numeric layout, but verify reader support first

The caution is portability. The spec may allow an encoding that your downstream reader doesn't fully implement.

How Parquet Delivers Speed Predicate Pushdown and Column Pruning

Parquet speed comes from not reading data you don't need. That sounds obvious, but it breaks into three distinct mechanisms with different failure modes.

Column pruning does the first cut

If your fact table has many columns and your SQL touches only a few, the engine can read only those column chunks. That's the most visible win in analytical systems. It's also why teams looking into SQL query optimisation often discover that file format and table layout matter as much as the query text.

In practical terms, a query like:

select order_date, region, revenue
from sales_facts
where order_date >= current_date - interval '7' day
select order_date, region, revenue
from sales_facts
where order_date >= current_date - interval '7' day
select order_date, region, revenue
from sales_facts
where order_date >= current_date - interval '7' day

doesn't need marketing attribution fields, shipping notes, or dozens of unused dimensions. With Parquet, the planner can often avoid those bytes entirely.

Row group skipping is where layout decisions start paying off

Parquet files store metadata that helps readers decide whether a row group is worth scanning. When the filter says region = 'EMEA', and a row group's statistics show only values outside that range, the engine can skip it.

Sort order matters. If you write rows in a random order, min and max stats become less selective. If you cluster related values together, those stats become much more useful.

Page-level behavior helps, but only when the data cooperates

Pages are the smallest practical units within a column chunk. Readers may avoid decoding parts of a chunk if metadata and filtering logic allow it. Dictionary-encoded pages can also make equality predicates cheaper because repeated values have already been normalized into dictionary references.

That said, these optimizations weaken when:

  • Free-text columns don't compress or prune cleanly

  • High-cardinality fields spread values across many groups

  • Point lookups need a tiny answer from a large file

  • Unsorted writes smear filter values across the dataset

Query Pattern

Columns Read

Row Groups Scanned

Bytes Read

Wall Time

Full table scan

Many

Most or all

High

Slowest

Column-pruned aggregation

Few

Most or all

Lower

Faster

Predicate-filtered analytical query

Few

Selected groups only

Lowest of the three

Fastest when stats are selective

The trap is assuming Parquet is “fast” in every sense. It's fast at planned, sequential analytical reading. It isn't built like an indexed serving store.

Parquet vs ORC vs Avro vs CSV and JSON

Choosing Parquet gets easier when you compare it to the formats teams usually debate.

Two analytics formats, one row format, two interchange formats

Parquet and ORC both target analytical scans. They're columnar, compressed, and designed for typed data. In broad practice, Parquet tends to win on ecosystem breadth across engines and lakehouse tools. ORC still makes sense in some Hive-centered environments.

Avro is a different animal. It's row-oriented and better suited to record-by-record writing, event interchange, and schema-rich pipelines where preserving row structure matters more than scan efficiency.

CSV and JSON are easier for humans and general-purpose tools, but they push parsing and typing work to read time. That usually makes them poor long-term choices for large analytical datasets.

Use Parquet when you expect repeated scans and selective column reads. Use Avro when records move through streaming systems. Keep CSV and JSON for interchange, debugging, or simple handoffs.

The real decision is workload fit

A lot of teams ask, “Which format is fastest?” The better question is, “Fast for what?”

  • Batch analytics: Parquet usually fits best.

  • Hive-heavy analytics stack: ORC may deserve a look.

  • Streaming records and event transport: Avro is often cleaner.

  • Ad hoc inspection and broad tool compatibility: CSV still wins.

  • Nested API payload exchange: JSON remains common despite the cost.

Format

Layout

Compression Efficiency

Scan Speed

Schema Evolution

Best Fit

Parquet

Columnar

Strong

Strong for analytics

Good, with caveats across readers

Data lakes, BI, batch analytics

ORC

Columnar

Strong

Strong for analytics

Good

Hive-oriented analytical environments

Avro

Row-oriented

Moderate

Better for row access than column scans

Strong for records

Streaming, message payloads, interchange

CSV

Row-oriented text

Weak

Weak on large analytical reads

None built in

Manual inspection, simple exports

JSON

Semi-structured text

Weak to moderate

Weak for large scans

Flexible but loosely enforced

API payloads, nested interchange

Partitioning File Sizing and Storage Best Practices

A common production story goes like this: the team chose Parquet, query times looked good in the first month, and then the lake filled with tiny files, uneven partitions, and tables that one engine scanned quickly while another struggled with planning overhead. The format did its job. The layout did not.

An infographic titled 4 Decisions That Shape Parquet Lake Performance, outlining key optimization strategies for data lakes.

Pick partition columns that people filter on

Partitioning works like a coarse table of contents. It helps readers skip entire folders before they even open Parquet footers. That is powerful, but only when the partition key matches real query patterns.

Date is the usual starting point because reporting and backfills often slice by day, week, or month. Region, environment, tenant tier, or business unit can also fit if those fields appear often in WHERE clauses. A bad key usually has one of two problems. It is either too broad to prune much, or so high-cardinality that it explodes into countless small directories and files. Partitioning by user ID is the classic mistake.

The trade-off matters more than many guides admit. Good partitioning speeds broad analytical reads. It does little for point lookups, because Parquet still needs footer metadata before it can find the right row group. If your workload includes frequent record-level fetches, no partition scheme will turn Parquet into a key-value store.

For teams building batch writers and compaction jobs, these python ETL orchestration tips are a practical companion because partitioning only holds up when the scheduler, retry logic, and writer outputs stay consistent over time.

File size determines whether object storage feels fast or clumsy

File sizing is where many lakes drift out of shape. Very small files increase listing overhead, metadata reads, and planning time. Very large files can reduce parallelism and make retries expensive. The sweet spot depends on your engine, network, and write path, but the goal is stable, scan-friendly files rather than whatever size happened to fall out of upstream micro-batches.

Row group sizing matters too. Larger row groups often improve sequential reads and compression, but they also make selective reads less precise if your sort order is poor. That is why file size and sort order should be chosen together, not as separate tuning knobs.

Object storage adds another layer of behavior. A data lake on S3-compatible storage does not behave like a local disk with cheap directory operations. The costs show up in list calls, open latency, and small-file fan-out. This primer on how S3 behaves as a file system is a useful reference for that mental model.

Watch a few signals in production: average file size by table, file count per partition, compaction lag, and query planning time versus scan time. If planning time keeps rising while data volume stays flat, file layout is often the first place to look.

Sorting and bucketing help, but only in the right places

Sorting is one of the highest-value habits for Parquet writers. When similar values land together, row group statistics become more selective, compression improves, and engines can skip more data with confidence. If analysts often filter on event_date, customer_region, or another selective field, sorting by that field can make footer metadata far more useful.

Bucketing solves a different problem. It can make distribution inside a partition more predictable, which helps some join-heavy pipelines. It also adds operational overhead and is not supported equally well across engines, so treat it as a workload-specific choice, not a default.

A practical checklist:

  • Partition for common filters: Start with fields that consistently appear in query predicates.

  • Keep file counts under control: Compact after streaming or highly parallel writes.

  • Sort within partitions: Better ordering makes row group stats worth trusting.

  • Track layout drift over time: Growing small-file counts and slower planning are early warnings.

  • Use bucketing selectively: Apply it where join behavior justifies the maintenance cost.

Schema Evolution Interoperability and the Quiet Pitfalls

Parquet's reputation for portability is deserved only up to a point. The uncomfortable production reality is that the spec moves faster than many reader stacks.

Open format doesn't mean uniform support

Apache's implementation status material shows support is fragmented by engine and release year, with some newer capabilities not yet broadly implemented and some features only partially supported across readers, according to the Parquet implementation status page. The versions guidance also warns that some features are forward-compatible while others are forward-incompatible.

That means older readers might still open a file but lose metadata, miss performance improvements, or fail on unsupported encodings. Recent work such as adaptive lossless floating-point encoding and new logical-type work widens the gap between what the format can express and what every engine in your stack can reliably consume.

The trap usually appears months after rollout

The dangerous failures aren't always loud. Sometimes one engine reads a field as expected, another returns nulls, and a third falls back to a slower decode path. You don't catch that in a quick smoke test if everyone validates only with the writer engine.

A few failure patterns show up repeatedly:

  • Logical type mismatches: Timestamps, decimals, and nested types may render differently across readers.

  • Encoding compatibility gaps: Newer encodings can surprise older engines.

  • Dictionary fallback behavior: Different writers may switch strategies differently, affecting file size and scan speed.

  • Footer-first latency: Even tiny reads still pay the metadata access cost discussed earlier.

If your team is managing Databricks-heavy pipelines and cross-engine consumers, a dedicated Databricks data quality approach becomes relevant because type drift and reader disagreement often show up first as quality incidents, not parser errors.

Feature

Spark 3.5

Trino 470

DuckDB 1.2

Athena

Core Parquet reading

Broadly supported

Broadly supported

Broadly supported

Broadly supported

Newer encodings

Verify per release

Verify per release

Verify per release

Verify per release

Nested logical types

Usually workable, test carefully

Test carefully

Test carefully

Test carefully

Metadata preservation

Varies by workflow

Varies by workflow

Varies by workflow

Varies by workflow

Don't ask only “Can this engine read Parquet?” Ask “Can this exact reader version safely read files from this exact writer configuration?”

Schema Drift Timeliness and Observability for Parquet Pipelines

A Parquet footer isn't just reader metadata. In a mature platform, it becomes an observability surface.

A five-step flowchart illustrating how Parquet file metadata is extracted, analyzed, monitored, and used for automated actions.

Schema drift usually announces itself in metadata first

When a writer adds a column, changes a type, drops a field, or alters sort behavior, those changes appear in file metadata before downstream dashboards fully reveal the blast radius. That makes footer inspection useful for early detection.

Common drift signals include:

  • Unexpected nullable fields: A new column appears but consumers don't populate it consistently.

  • Type widening: Integer-like fields become broader numeric types and downstream logic changes behavior.

  • Dropped or renamed fields: Readers may still succeed, but business logic starts reading emptier results.

  • Partition mismatch: Directory naming and file schema stop telling the same story.

Timeliness and file semantics are linked

Parquet pipelines also create distinctive freshness problems. A table may look present in storage while a subset of files is still incomplete, delayed, or written with inconsistent metadata. Streaming writers can leave partial outputs that confuse consumers if your platform marks the dataset “ready” too early.

That's why engineers monitor more than arrival time. They also inspect file counts, modification timing, schema consistency across partitions, and writer metadata such as created_by and sorting information.

One way to operationalize that is with data lake monitoring that watches file-level and dataset-level signals together instead of treating storage as a black box. In that context, digna fits as one option for teams that want in-environment monitoring of schema changes, timeliness, anomalies, and validation across lake and warehouse pipelines.

What to watch in production

If I were onboarding a new analytics engineer to a Parquet-heavy lake, I'd tell them to start with four checks:

  1. Footer diffs between daily loads: Catch silent schema changes early.

  2. Small-file growth by partition: Query pain often starts here.

  3. Reader-writer compatibility matrix: Keep an explicit list by engine version.

  4. Timeliness tied to complete dataset readiness: Don't alert only on missing folders.

The most useful Parquet metadata isn't there to make the file elegant. It's there to help you detect trouble before users ask why the dashboard looks wrong.

Parquet works well when your workload matches its design, but it needs disciplined monitoring once multiple engines, writers, and downstream consumers enter the picture. digna helps teams track schema change, timeliness, validation, and data behavior inside their own environment, which is exactly where Parquet pipeline issues tend to surface first. If your lakehouse runs on Parquet and your incidents start with silent drift instead of loud failures, it's worth a look.

Footer metadata makes schema drift visible early, but someone still has to watch it — that is the job of a data quality management layer sitting over the lake.

Frequently asked questions

Why did Parquet become the default columnar format?

Parquet became the connective tissue across Spark pipelines, warehouse exports, object-storage lakes and table formats like Iceberg, Delta Lake and Hudi. Teams adopted it because it is open, typed and broadly supported, and because it separates physical storage from engine choice — a file written by one tool stays readable by another.

Which encodings and compression codecs does Parquet support?

Parquet supports page-level encodings including PLAIN, RLE, DELTA_BINARY_PACKED, RLE_DICTIONARY and BYTE_STREAM_SPLIT, plus codecs such as SNAPPY, GZIP, BROTLI, ZSTD and LZ4_RAW. Encoding restructures values inside a page; compression then shrinks the encoded byte stream. Pick the codec by what you are optimising — read speed, storage cost or write throughput.

What is predicate pushdown in Parquet?

Predicate pushdown lets a reader skip data using metadata before decoding any of it. Row-group statistics record the range of values present, so a filter on region = 'EMEA' can discard row groups whose statistics prove no match. It works alongside column pruning, which removes unused columns first.

How should Parquet files be partitioned and sized?

Partition on columns people actually filter on — partitioning acts as a coarse table of contents that lets readers skip whole folders before opening any footer. On sizing, very small files inflate listing and planning overhead while very large ones reduce parallelism and make retries expensive. Sorting by a selective field sharpens row-group statistics.

Is Parquet schema evolution safe across engines?

Only up to a point. Apache's implementation status material shows support fragmented by engine and release year, with some newer features only partially implemented. The dangerous failures are quiet ones: one engine reads a field correctly while another returns nulls. Validating only with the writer engine hides that for months.

✦ Generato con l'intelligenza artificiale

Condividete su X
Condividete su X
Condividete su Facebook
Condividete su Facebook
Condividete su LinkedIn
Condividete su LinkedIn

Il team dietro la piattaforma

Un team con sede a Vienna di esperti di AI, dati e software, supportato

da rigore accademico ed esperienza enterprise.

Il team dietro la piattaforma

Un team di esperti di IA, dati e software con sede a Vienna, forte di rigore accademico ed esperienza aziendale.

Prodotto

Integrazioni

Risorse

Azienda

INDEXED BYIndexerNow INDEXED BYIndexerNow