• new

    The major Release 2026 is live - Bringing Data Observability Into Your Code

  • new

    Contribute to the Future of AI & Data Innovation

  • new

    • Release 2026.06 - Bringing Data Observability Into Your Code

  • new

    • Contribute to the Future of AI & Data Innovation

How Do You Validate Data? a Guide for Modern Pipelines

|

8

min read

You're probably asking this because something already broke. A dashboard number jumped overnight, a model started producing odd outputs, or a reconciliation failed and nobody trusts the pipeline until someone digs through logs. That's the core context behind the question how do you validate data. It isn't academic. It's operational.

Good validation doesn't mean adding a few null checks and calling it done. It means deciding what must be true at the record level, what must stay stable across the pipeline, what can drift safely, and what should stop the line immediately. In modern stacks, that also means guarding against problems that basic guides ignore, especially silent schema drift and the cost of over-enforcing rules that don't matter to the business.

Table of Contents

Why Data Validation Is More Than Just Checking Boxes

A revenue dashboard rarely fails because one engineer forgot a check. It fails because teams treat validation as a one-time cleanup step instead of part of pipeline design. By the time someone notices the numbers are wrong, bad data has already moved through transformations, joins, BI models, and downstream reports.

That's why data validation belongs at multiple checkpoints, not just at the final output. The strongest pattern is prevention first. Validation rules should catch problems where data enters the system, and then again after transformation steps where logic can introduce new errors. The Great Expectations overview of pipeline validation makes that point clearly by emphasizing validation at ingestion and after each transformation, along with repeatable documentation for auditability.

Trust is the real output

Organizations say they want clean data. What they really want is trustworthy decisions. If finance has to question every report, if analysts have to manually inspect every refresh, or if ML engineers don't trust training inputs, the platform is technically running but operationally failing.

Validation protects trust in a few concrete ways:

  • It blocks known bad inputs: Required fields, valid formats, and business rules stop obvious defects early.

  • It narrows incident scope: When checks run at each hotspot, teams know where the issue entered.

  • It makes failures actionable: A failed rule is more useful than a vague “dashboard looks wrong” complaint.

  • It supports compliance: Repeatable validation and published metadata give teams an auditable process.

Practical rule: Don't validate data based only on what you've seen before. Validate it against what the business says must be true.

Reactive cleaning is too late

A lot of teams still rely on downstream cleanup. That sounds practical until a broken field has already been aggregated into reports or pushed into customer-facing systems. Cleaning after the fact is slower, more expensive, and harder to audit.

A better operating model looks like this:

  1. Define business expectations before code.

  2. Enforce them at ingestion.

  3. Re-check them after every meaningful transformation.

  4. Log results so failures are traceable and repeatable.

Validation isn't bureaucracy. It's the engineering control that keeps bad records, stale data, and broken schemas from becoming business decisions.

The Core Dimensions of Data Validation

A pipeline can pass every basic check and still damage trust. The common failure is not a null in a required field. It is data that looks acceptable, moves through the stack, and stops matching business reality, undetected, after a source change, a delayed load, or a new code path upstream.

That is why validation needs more structure than a pass or fail gate. Teams need a way to decide what must be enforced, what should be monitored, and what can be tolerated briefly because the cost of blocking is higher than the cost of review. A useful starting point is six quality dimensions. They give teams a shared vocabulary for deciding where risk sits and how strict each control should be.

An infographic titled The Core Dimensions of Data Validation listing six key factors for achieving good data.

Good data has six dimensions

These dimensions sound familiar because they are. The mistake is treating them as a checklist instead of separate failure classes with different business costs.

Dimension

What it means in practice

Typical impact when it fails

Accuracy

The value reflects the real-world event or entity

Wrong pricing, wrong balances, wrong customer status

Completeness

Required data is present where the process depends on it

Broken joins, unusable reports, missing model features

Consistency

The same concept is represented the same way across systems

Reconciliation issues, duplicate logic, reporting disputes

Timeliness

Data arrives within the window the business expects

Stale dashboards, delayed operations, missed SLAs

Uniqueness

A record or entity appears once when it should

Duplicate customers, duplicate charges, inflated counts

Validity

Values conform to allowed formats, domains, and rules

Parse failures, rejected events, invalid transactions

Timeliness deserves more attention than it usually gets. I have seen teams approve a dataset because every field passed type and range checks, while operations were working off yesterday's data. The table was valid. The decision was still wrong.

The same applies to uniqueness and consistency. A duplicate transaction can be more expensive than a missing optional field. A status code that means one thing in the source system and another thing in the warehouse can pass schema validation and still break finance reporting. Validation works better when severity follows business impact, not technical neatness.

Different validation types catch different classes of risk

Each dimension needs a different kind of control. If teams only validate structure, they miss meaning. If they only watch distributions, they miss hard rule violations. Good coverage comes from combining several types of checks and assigning each one an enforcement mode.

Use this mapping:

  • Schema validation checks structure. Column presence, data type, nullability, and contract changes.

  • Syntactic validation checks format. Dates, currency codes, identifiers, booleans, and standardized text patterns.

  • Semantic validation checks business meaning. End date after start date, status transitions that match the workflow, prices that align with product rules.

  • Relational validation checks cross-table integrity. Foreign keys, orphaned records, and parent-child completeness.

  • Statistical validation checks behavioral drift. Volume shifts, null-rate changes, cardinality changes, and distribution anomalies.

Silent schema drift sits between these categories and causes some of the hardest incidents to diagnose. A source team can widen a field, repurpose an enum, change timezone handling, or start sending a new optional column that downstream code ignores. Nothing crashes. The numbers just stop lining up. In these situations, automated contract checks and profile-based monitors earn their keep. Teams evaluating free data validation tools for schema checks and drift monitoring should look for both hard-rule enforcement and trend-based alerts, because one without the other leaves blind spots.

A field can be valid by type and still be invalid for the decision it supports.

That is the gap basic guides often miss.

A practical rule design process starts in plain language. Define what must be true, who depends on it, what happens if it fails, and whether the pipeline should block, quarantine, warn, or log the event for review. IBM's overview of key data quality dimensions is useful here because it frames quality as fitness for use, which is exactly how validation should be prioritized in production systems.

The final design choice is economic. Blocking every anomaly sounds disciplined, but it can stop revenue operations, delay downstream consumers, and flood teams with low-value alerts. Letting everything pass is worse. Strong validation programs separate high-risk failures from tolerable variation, enforce the former automatically, and monitor the latter with clear ownership. That is how validation improves reliability without turning the pipeline into a constant incident generator.

Implementing Validation at Record and Pipeline Levels

Validation works best when you apply it at two levels at once. First, inspect each record for rule violations. Then inspect the pipeline as a system for drops, mismatches, and structural breaks. If you only do one, gaps stay open.

Start with row-level rules

At the record level, the baseline is straightforward. Industry best practices mandate row-level validation applying eight specific rules, required fields, type checking, format validation, range constraints, uniqueness, referential integrity, business logic, and cross-field validation, to every record, coupled with audit logging that tracks pass/fail counts per run for compliance and debugging, as described in Flatfile's guide to data validation.

A practical SQL pattern looks like this:

select
  order_id,
  customer_id,
  order_date,
  amount,
  case when order_id is null then 'fail_required_order_id' end as required_check,
  case when amount < 0 then 'fail_amount_range' end as range_check,
  case when order_date > current_date then 'fail_future_order_date' end as business_rule_check
from raw.orders;
select
  order_id,
  customer_id,
  order_date,
  amount,
  case when order_id is null then 'fail_required_order_id' end as required_check,
  case when amount < 0 then 'fail_amount_range' end as range_check,
  case when order_date > current_date then 'fail_future_order_date' end as business_rule_check
from raw.orders;
select
  order_id,
  customer_id,
  order_date,
  amount,
  case when order_id is null then 'fail_required_order_id' end as required_check,
  case when amount < 0 then 'fail_amount_range' end as range_check,
  case when order_date > current_date then 'fail_future_order_date' end as business_rule_check
from raw.orders;

That isn't glamorous, but it's effective. The point is to make each failure explicit and classifiable.

You'll usually need a mix of checks:

  • Required fields: Block nulls in keys, dates, and operationally critical attributes.

  • Type and format checks: Enforce dates as dates, Booleans as Booleans, codes against reference formats.

  • Range constraints: Catch impossible or unsafe values before they distort downstream logic.

  • Cross-field rules: Validate relationships such as start and end dates, debit and credit signs, or country and postal format combinations.

Screenshot from https://digna.ai

Then validate the pipeline as a system

Row checks won't catch everything. A pipeline can pass record-level rules and still be broken if half the data never arrived, if a join exploded duplicate rows, or if source and target counts no longer reconcile.

That's where pipeline-level checks matter:

  • Row count comparison: Compare source and target counts after loads and transformations.

  • Duplicate monitoring: Check whether expected unique keys remain unique after joins or unions.

  • Referential integrity: Confirm foreign key values exist in referenced tables.

  • Aggregate sanity checks: Compare totals, distributions, and null patterns across stages.

  • Migration parity checks: For critical fields during migrations, validate every row when the business can't tolerate silent loss.

The Digna migration validation guidance is especially useful here. It notes that for critical fields during data migrations, 100% row-level validation is necessary to confirm record counts match between source and target while verifying no duplicates were created, no records were dropped unnoticed, and no partial records exist, while large datasets can use statistically significant sampling with anomaly detection when full manual checking isn't practical.

Keep heavy checks in the warehouse

Large-table validation often fails because teams pull too much data out of the warehouse and inspect it in application code. That's slow, expensive, and hard to scale. Push the heavy work to the database whenever possible.

In-database validation is a better fit for:

  • uniqueness checks on compound business keys

  • referential integrity across datasets

  • threshold and range checks on large tables

  • profile comparisons on null rates or category distributions

If you want a lightweight starting point before building your own framework, these free data validation tools from digna are worth reviewing alongside options like dbt tests, Great Expectations, and warehouse-native SQL checks.

The warehouse is already optimized to scan, compare, aggregate, and join. Let it validate there instead of exporting the problem somewhere else.

Automating Validation and Managing Failures

Manual spot checks are fine for debugging. They aren't an operating model. If validation doesn't run automatically whenever data changes, your team is depending on luck and curiosity.

A diagram illustrating an automated data validation workflow, from ingestion and validation to cleaning and reporting.

Automate checks where data changes

The cleanest automation pattern is event-driven or orchestration-driven. Run checks after ingestion, after major transformations, and before publishing data to serving layers. Airflow, dbt, and warehouse tasks all support this pattern.

A durable sequence looks like this:

  1. Ingest data and validate schema immediately.

  2. Run record-level rules on landing tables.

  3. Run aggregate and parity checks after transforms.

  4. Write pass/fail results to an audit table.

  5. Trigger the appropriate failure action.

Observability and validation start to converge. Validation tells you whether a rule passed. Observability helps you understand trend changes, timeliness gaps, and whether the same failure is recurring across runs. A useful primer is this overview of data observability concepts and workflow design.

A short demo can help make the orchestration pattern concrete:

Choose failure actions by business risk

Not every failed check deserves the same response. At this juncture, teams often either overreact or underreact.

A useful decision model is to classify failures into three buckets:

Failure type

Example

Best action

Hard stop

Missing primary key, invalid financial period, referential break in critical fact table

Halt the pipeline

Quarantine

A subset of rows violates format or business logic

Route bad records for review

Warn and continue

Minor drift in a noncritical descriptive field

Alert and monitor

The trade-off is economic, not just technical. Industry data shows that data quality issues cause 20-30% of revenue loss in finance and healthcare, yet no standard framework exists to calculate the breakpoint where validation effort outweighs the marginal risk reduction, according to Twilio's discussion of validation techniques and business impact. That gap matters. Teams need to decide where strict enforcement pays for itself and where it creates friction without much risk reduction.

If your stack also feeds generative systems, data quality and model monitoring start to overlap. When you're evaluating downstream controls, it helps to find the right LLM monitoring platform so you can see how input reliability, prompt behavior, and production monitoring fit together.

Avoid alert fatigue

Too many teams flood Slack and email with low-value alerts until nobody reads them. Good automation creates signals, not noise.

A few habits work well:

  • Route by severity: Paging should be rare. Most issues belong in dashboards or ticket queues.

  • Group related failures: If ten tables fail because one upstream schema changed, send one incident.

  • Include context: Every alert should say what failed, where, when, and what the system did next.

  • Review patterns periodically: The Cube guide to validation best practices describes a useful cadence in enterprise settings where teams review error patterns quarterly and update rules annually, rather than leaving stale thresholds in place.

Alerts should tell an engineer what happened and what to do next. If they only announce “validation failed,” they're unfinished work.

Advanced Strategies for Enterprise Scale

At enterprise scale, static rules still matter, but they stop being enough. You need controls for changes nobody explicitly coded for, especially when BI tools, ETL services, and source applications keep evolving underneath you.

Silent schema drift is the enterprise problem basic guides miss

A lot of failures don't come from a null or an out-of-range value. They come from a subtle structural change. A column gets renamed, a type changes from integer to string, or an upstream tool auto-adjusts a schema in a way that keeps ingestion running but breaks downstream logic.

That's why schema tracking should be treated as validation, not just metadata hygiene. The risk is larger than many teams assume. Recent 2025-2026 industry reports indicate that 65% of ML pipeline failures stem from undetected schema changes rather than data value errors, yet standard validation protocols rarely include schema version tracking, as cited in this discussion of schema drift and pipeline failures.

Screenshot from https://digna.ai

A mature schema validation practice includes:

  • Version tracking: Record structural changes over time.

  • Compatibility checks: Decide which changes are backward compatible and which should block publication.

  • Ownership: Make one team responsible for approving schema changes.

  • Downstream impact checks: Link schema change alerts to affected models, dashboards, or APIs.

Add anomaly detection and timeliness

Hard-coded rules only catch what you already know to look for. Enterprise systems need another layer that detects unexpected change in distributions, null patterns, category mixes, and delivery timing.

This is one place where platform support helps. Tools such as dbt tests and Great Expectations handle rule-based validation well. For in-database monitoring of anomalies, timeliness, record-level rules, and schema changes, one option is digna, which runs analyses inside the customer environment and surfaces trends, delays, and structural shifts without exporting production data.

Timeliness deserves equal treatment with content checks. Late data can still invalidate dashboards while every row passes type and format rules. Monitor expected arrival windows and flag delayed loads before stakeholders find stale reports on their own.

Teams building AI support workflows run into a similar issue at the application layer. Data can be structurally valid but still lead to unreliable outputs if inputs and instructions aren't controlled. That's why guidance on prompt design for reliable support AI is useful in parallel with data validation. It addresses another version of the same discipline: defining acceptable inputs and reducing silent failure modes.

Governance keeps rules from decaying

Validation rules age badly when nobody owns them. One team writes them, another team changes the source process, and six months later the checks are either noisy or irrelevant.

A sustainable model includes:

  • Plain-language rule definitions: Business stakeholders should be able to review them without reading SQL.

  • Assigned owners: A data owner, steward, or governance team should maintain each rule.

  • Audit logging: Keep pass/fail counts and run metadata.

  • Scheduled review: Revisit thresholds, assumptions, and exceptions regularly.

That governance layer is what turns validation from a script collection into an operational control system.

A Pragmatic Framework for Data Validation

A practical validation program starts with risk, not coverage. Teams get into trouble when they write dozens of checks for low-impact data, then miss the few failure modes that can corrupt revenue reporting, break customer workflows, or push bad features into models. A better framework asks two questions first: what can fail undetected, and what is the business cost if it does?

A five-step framework infographic illustrating the pragmatic process for ensuring effective and reliable data validation practices.

A working model for teams that need results

Start with the data assets that carry the highest operational or regulatory risk. In practice, that usually includes financial measures, customer identifiers, compliance fields, and model inputs. Define a small rule set for each one, then tie every rule to an action. If a check fails, decide whether the pipeline should stop, quarantine the affected data, or continue with an alert.

Use a mixed enforcement model because not every defect should be treated the same way:

  • Apply strict rules to invariants: Primary keys, required fields, referential integrity, and fixed formats.

  • Use anomaly detection for changes that are hard to enumerate in advance: Distribution shifts, null spikes, volume changes, and delayed arrivals.

  • Track schema drift explicitly: Silent column changes and type changes often break downstream logic before anyone notices.

  • Assign an owner to every rule: Someone has to review noise, approve exceptions, and retire outdated checks.

Many basic guides fall short in this regard. They treat validation as a pass or fail gate. Production systems need a risk-based model instead. A missing primary key in a finance table deserves a hard stop. A minor shift in a low-priority attribute may only need a warning and a ticket. The point is not maximum enforcement. The point is reliable data at a cost the team can sustain.

What to do first this week

Start with one disputed pipeline. Pick the one people already question in meetings, because it already has visible business impact and a natural feedback loop.

Then do five things:

  1. Name the five business conditions that must hold true.

  2. Map each condition to a technical check at the record or pipeline level.

  3. Classify each failure by impact: stop, quarantine, warn, or log only.

  4. Add run history so the team can see repeated failures and drift over time.

  5. Review false positives after the first week and tighten the noisy rules.

That sequence works because it forces trade-offs early. Teams learn which controls protect trust and which ones only create alert fatigue. It also exposes actual failure modes before anyone invests in a large rule catalog that will be expensive to maintain.

Data validation works best as an operating system for reliability. It combines deterministic rules, drift detection, schema awareness, ownership, and failure handling into one process that can survive changing sources and growing pipeline complexity.

If you want a practical way to combine in-database validation, anomaly detection, timeliness monitoring, and schema tracking in one workflow, take a look at digna. It's built for teams that need to validate records, catch drift, and monitor pipeline reliability without moving production data out of their own environment.

Because silent schema drift slips past record-level rules, digna Schema Tracker covers the version-tracking side of validation by detecting added, removed and retyped columns before downstream logic breaks.

Frequently asked questions

How do you validate data in a pipeline?

Validate at multiple checkpoints rather than only at the output. The article's sequence is: validate schema at ingestion, run record-level rules on landing tables, run aggregate and parity checks after transforms, write pass/fail results to an audit table, then trigger the failure action that matches the business risk.

What are the six dimensions of data quality?

The six dimensions are accuracy, completeness, consistency, timeliness, uniqueness and validity. The article treats them as separate failure classes with different business costs: a duplicate transaction can cost more than a missing optional field, and a table can pass every type check while operations still work from yesterday's data.

What row-level validation rules should every record pass?

Flatfile's guide, cited in the article, lists eight: required fields, type checking, format validation, range constraints, uniqueness, referential integrity, business logic and cross-field validation. Pair them with audit logging of pass/fail counts per run. A simple SQL CASE pattern can flag null order IDs, negative amounts or future order dates.

What should happen when a data validation check fails?

Match the response to business risk. Hard stops halt the pipeline for issues like a missing primary key or an invalid financial period. Quarantine routes a subset of bad rows for review. Warn-and-continue suits minor drift in a noncritical descriptive field, which gets an alert and monitoring instead of blocking.

Why is schema drift a data validation problem?

Structural changes often slip past value checks: a renamed column or an integer-to-string type change can keep ingestion running while breaking downstream logic. Reports cited in the article attribute 65% of ML pipeline failures to undetected schema changes, so version tracking, compatibility checks and ownership belong in validation.

✦ Generated with Artifical Intelligence

Share on X
Share on X
Share on Facebook
Share on Facebook
Share on LinkedIn
Share on LinkedIn

Meet the Team Behind the Platform

A Vienna-based team of AI, data, and software experts backed

by academic rigor and enterprise experience.

Meet the Team Behind the Platform

A Vienna-based team of AI, data, and software experts backed by academic rigor and enterprise experience.

Product

Integrations

Resources

Company

INDEXED BYIndexerNow INDEXED BYIndexerNow