Can Data Fix Itself? Understanding Autonomous Data Quality
|
6
min di lettura

Everyone who works with data has had the call. A number in a report is wrong, and somebody needs to find out why before the meeting starts. The investigation that follows is almost always the same: which feed, which load, which column, how far has it spread, and what do we do about it?
For twenty years the tools have got steadily better at one part of that job: noticing. Rules catch what somebody anticipated. Observability catches what moved. But between the alert and the fix there is still a person, a queue and, very often, a Monday morning.
This article asks whether that has to stay true. Can data fix itself? The honest answer is: partly, under conditions you can write down. And the conditions turn out to be more interesting than the automation.
Data Quality and Data Observability: Two Disciplines, Neither Enough
Data quality is the degree to which data is fit for the purpose it is used for, measured against an expectation somebody had to write down. It is a property of the data itself: values, keys, relationships. It is relative to a purpose, because a departure time good enough for a monthly report is not good enough for filing an airport slot. And it is declarative. A rule only ever finds what somebody already thought of.
Data observability is the ability to understand the health and behaviour of your data from the signals it emits, without knowing in advance what will go wrong. Freshness, volume, schema, distribution and lineage are learned from history rather than specified by the business, which is why observability scales to every table. It is also its limit: observability can tell you that something changed, never that something is wrong.
Neither discipline contains the other. Consider a block time of 92 minutes on a route that normally takes 148. Row counts are normal, the table is fresh, the schema has not changed, and the value sits comfortably inside any global range. It passes every check and it is still wrong. This is the gap described in our article on data that looks wrong but passes your rules, and the full comparison of the two disciplines is in Data Observability vs Data Quality.
What Is Autonomous Data Quality?
Autonomous data quality is the continuous, mostly AI-driven ability to detect, investigate, assess and remediate data quality issues with minimal human intervention. Four verbs, and the last three are what separate it from the tooling most teams already have:
Detect: something moved far enough to be worth attention.
Investigate: which feed, which load, which column.
Assess: how bad it is, and what depends on it.
Remediate: quarantine, hold or re-run, and record exactly what was done.
Almost every product on the market today stops after the first verb. Like observability, autonomous data quality adapts: thresholds and priorities follow the data instead of freezing on the day somebody wrote them. In practice it is the attempt to use the signals observability produces to generate the quality work that rules used to require.
Why This Is Becoming Possible Now
Two things are changing at the same time. The first is the shape of the data platform. The central warehouse, owned by one team that agreed what a customer or a passenger was, is giving way to data products with their own owners, contracts and release cycles. Quality now has to hold at every boundary the data crosses, not just where it lands, and nobody owns the whole chain any more.
The second is the arrival of agents on top of that platform. An agent does not need to be told how. It needs to be told what, and how far it may go. Take a routine reload. Today an engineer finds the stored procedure and its parameters, the source API with its authentication and paging, writes the retry and error handling, and rewrites all of it when a partner changes their export. With an agent the instruction becomes a sentence: reload yesterday's boarding records for one flight. The agent finds the procedure and the API in the catalog, sequences the calls, retries, and verifies the row count itself.
That is a different integration model from anything we have built before. We will spend less time specifying how a reload happens, and more time specifying what may be reloaded, by whom, and without asking.
One Flight Leg, Five Different Answers
To make this concrete, take Alpenwing Airways, a fictional mid-size European carrier whose problems are entirely real. Seven systems describe the same flight: bookings, departure control, the airport operational database, maintenance, crew rostering, revenue accounting and eleven codeshare partner feeds. None of them is wrong. They were built for different jobs at different times, and the warehouse is where their disagreements become visible. Here is what an agent does with five alerts on a single flight leg, WG 402 from Vienna to Warsaw.
1. A partner feed switches from IATA to ICAO airport codes overnight
The average length of the departure airport column moves from 3.00 to 4.00 on one source, and 1,284 rows fail the allowed-values check. Nothing else on that feed moved. A Data Steward approved the mapping LOWW to VIE months ago, the published ICAO register confirms it, and the fix is reversible. This is the easiest case there is, and it is close to what is possible today.
2. The flight boards 36 passengers onto a 180-seat aircraft
A load factor of 0.20, on a route that has not dropped below 0.72 in two years, while every other leg that day looks normal. There is an approved pattern for exactly this: reload the leg, then verify it against a second, independent count from the booking system. Lineage shows three marts and the daily operations report depend on it, so the agent holds them until the count is confirmed.
3. Departure control counts 168 passengers, the airport database counts 171
A scheduled reconciliation fails by three passengers, and observability sees nothing at all: 168 is a normal number and so is 171. The difference turns out to be definitional. Is an infant travelling on a lap a passenger? No catalog entry answers that and no approved pattern covers it. The agent can do the entire analysis. The definition stays a human decision.
4. The airport operational feed is three hours late
The feed normally lands before 02:30. At 05:40 it still has not arrived, and the operations report is due at 06:00. No data quality rule fires, because the data is not wrong, it is simply absent. Deferring dependent loads on a late arrival is an approved pattern, so the agent uses lineage and the load log to hold exactly the reports that would otherwise run on yesterday's figures, and nothing else.
5. Destination city arrives as free text
Wiedeń, Viena and Wien all mean Vienna. Distinct values jump from 41 to 63 in a day, and 2,140 rows match no reference list. The open internet is useful for learning that Wiedeń is Polish for Vienna, but it is never sufficient on its own. A mapping is only safe against a declared domain: the 94 airports Alpenwing actually serves. This is where the technology is heading rather than where it is today.
The Four Checks That Decide Whether an Agent May Act Alone
Every one of those scenarios passes through the same gate, built from four checkable questions. None of them is the model's own confidence, which is not calibrated and should never be the reason a warehouse is changed.
Has this exact pattern been approved before? If a Steward approved it once, the tenth occurrence is a lookup rather than a judgement. It is the strongest single signal, but it is neither required nor sufficient.
Is the action reversible? Quarantine, hold, re-request and reload can all be undone. Overwriting a value in place cannot. Reversible actions earn far more freedom.
Is the source authoritative and dated? A published register with a publication date justifies acting. A plausible answer with no citation does not, however confident it sounds.
How large is the blast radius? One quarantined row is not the same as a dimension every mart joins to, which in turn is not the same as a figure already filed with a regulator.
The order in which an agent consults its sources matters just as much. Your own knowledge base comes first: the catalog, lineage, contracts and the library of approved patterns. Published registers come second. The open internet is for orientation only, and the model's own memory is the weakest source of all, because an answer nobody looked up is indistinguishable from one that was researched. An agent that cannot name its source does not get to act.
How much freedom the agent gets is set per domain and per data product, never once for the whole company. A regulated bank may allow proposals only, each waiting for a named approver. An airline's operations data may let reversible actions run alone while anything irreversible is proposed. A marketing data product may run most fixes unattended and be reviewed weekly in aggregate. And whatever the setting, every action carries its evidence, every action can be rolled back on its own, and approved patterns are monitored and expire, because the world can move underneath a pattern that used to be right.
What Autonomous Data Quality Still Cannot Do
The limits have moved. They have not disappeared, and three of them are permanent.
It cannot invent a source of truth. Whether a passenger actually boarded is recorded by the boarding scan and nowhere else. If both copies of a figure are wrong together, reconciliation will happily confirm them, because it compares rather than verifies.
It cannot define the purpose. Whether an infant counts as a passenger, which system is authoritative, and whether history may be restated after a figure was reported to a regulator are business decisions. Somebody has to own them, in writing.
It does not fix the upstream. A feed that is silently repaired forever is a feed that is never repaired at the source, because the supplier never feels the pain. That is why every automatic fix must still be reported.
What Happens to the Data Steward
Every data team knows the Monday morning: forty-one alerts in the queue from the weekend, three of which are real problems. Autonomous data quality does not make that pile disappear. It removes the repeatable part of it.
What goes away is reading the same alert for the fortieth time, chasing a partner about a feed that broke again, and re-running loads by hand at seven in the morning. What replaces it is curating the library of approved resolution patterns, deciding what an ambiguous value actually means, auditing the agent's decisions in aggregate, and setting how much freedom each data product gets. The Data Steward stops fixing individual problems and starts deciding how problems get fixed. That is a more senior job than most stewards hold today, closer to policy than to operations. Nobody is removed. The work that scaled badly is.
Final Thought: Precedent, Reversibility and Evidence, Not Confidence
A general-purpose agent with database credentials is not autonomous data quality. Without a catalog it does not know whether a column holds a code or a label, and it guesses with certainty. Without lineage it cannot scope a re-run or judge a blast radius. Without a pattern library every occurrence is the first one, nothing compounds, and the Steward never gets any time back. Without those three, a language model should not be near your warehouse.
That is why autonomous data quality belongs inside the data quality platform rather than beside it. The foundation already has to be there: digna Data Anomalies learns what normal looks like without manual thresholds, digna Data Validation enforces auditable record-level rules, digna Timeliness learns arrival patterns alongside the schedules you declare, digna Schema Tracker catches structural change, and digna Data Analytics tracks how the metrics themselves trend over time. All of it runs in-database, without data leaving your environment. The trust gradient, the autonomy gate, the evidence trail and the pattern library described here are the direction digna is building toward.
Can data fix itself? Not on its own. But with precedent, reversibility and evidence in place, a growing share of it can be fixed without waiting for Monday.
See the foundation autonomous data quality is built on.
digna learns what normal looks like across every table, in-database, cloud or on-premise, with no manual thresholds to maintain. It is the layer of evidence an agent needs before it can be trusted to act.
Book a Personalised Demo → Explore the digna Platform
Frequently asked questions
What is autonomous data quality?
Autonomous data quality is the continuous, mostly AI-driven ability to detect, investigate, assess and remediate data quality issues with minimal human intervention. Most tools today stop at detection; the autonomous part is the investigation, the impact assessment and a recorded, reversible fix.
How is autonomous data quality different from data observability?
Data observability learns baselines for freshness, volume, schema and distribution and tells you that something changed. Autonomous data quality uses those signals as triggers, then investigates the cause, assesses what depends on the data and takes a bounded action such as a quarantine, hold or reload.
When is an AI agent allowed to fix data on its own?
Four checks decide it: whether the exact pattern was approved before, whether the action is reversible, whether the source is authoritative and dated, and how large the blast radius is. The model's own confidence is never one of the checks, because it is not calibrated.
Which sources should a data quality agent trust?
Its own knowledge base first: the catalog, lineage, contracts and approved patterns. Published registers such as IATA or ICAO code lists come second. The open internet is useful only for orientation, and the model's unsourced memory is the weakest source. An agent that cannot name its source should not act.
Will autonomous data quality replace the Data Steward?
No. It removes the repetitive part of the job, such as rereading the same alert or re-running loads by hand. The Steward moves to curating approved resolution patterns, deciding what ambiguous values mean and setting how much freedom each data product gets, which is closer to policy than operations.



