• nouveau

    La grande Release 2026 est disponible – Intégrez la Data Observability au cœur de votre code

  • nouveau

    Contribuez à l'avenir de l'innovation en matière d'IA et de données

  • nouveau

    • Release 2026.06 - Intégrer la Data Observability au cœur de votre code

  • nouveau

    • Contribuez à l'avenir de l'innovation en matière d'IA et de données

Data Quality Control in Research: A Practical Guide

|

0

minute de lecture

You usually don't find data quality problems when the dataset is still warm from collection. You find them after a spreadsheet has been copied, a file has been merged, a field name has changed, and someone asks why one row now shows an impossible date or a missing follow-up that nobody flagged in time. That's the shape of data quality control in research, it's not a tidy checklist at the end, it's a chain of controls that has to survive every handoff.

Table of Contents

Why Data Quality Control in Research Fails Quietly

A research team can spend weeks collecting clean-looking records and still inherit a mess. The failures often stay hidden because transcription, transfer, updates, and saves all create new chances for silent drift, and many teams only inspect quality at collection time. By the time someone notices an outlier or a missing variable, the source form may be buried, the schema may have moved, and the cleanup becomes reconstruction.

That's why the strongest systems treat quality as a lifecycle obligation, not a one-time audit. A good starting point is to look at the structural reasons projects break and fix the process around them, not just the symptoms, as outlined in this practical note on why data quality projects fail.

Practical rule: if a record can change hands, it can lose integrity.

The rest of this guide stays grounded in five things that hold up under pressure, validation at each handling step, measurable thresholds, provenance and documentation, reproducibility, and audit trails. That mix matters because the biggest errors aren't dramatic, they're ordinary, and ordinary errors are exactly what slip past teams who only check at the finish line.

The Five Dimensions of Data Quality

A diagram illustrating five core pillars of data quality control: completeness, correctness, concordance, plausibility, and currency.

Good quality control starts with shared language. If a team says “the data look fine” without naming what they checked, they're usually talking past each other. The five dimensions, completeness, correctness, concordance, plausibility, and currency, give the team a way to separate missingness from contradiction, and stale data from impossible values.

What each dimension catches

Completeness is about gaps. A missing follow-up visit, an empty lab field, or a skipped consent date all live here, and they aren't the same problem as wrong values. A simple weekly completeness report on key fields is enough to surface where the holes cluster.

Correctness asks whether a value matches the source. If a birth date is typed wrong or a lab result is copied from the wrong line, the number may be perfectly formatted and still be false. A practical check is a targeted source-document comparison on a small sample of records that matter most.

Concordance is cross-system agreement. Two databases that disagree on the same participant's enrollment status create reconciliation work later, so compare identifiers, visit dates, and status flags across systems early. At this point, many teams overtrust the “master” table and miss the fact that the master is just the last place the error landed.

Plausibility catches values that are technically valid but not believable. An age of 0 in an adult cohort, a discharge date before admission, or an implausible timestamp sequence should trigger review even if the field passes format checks. This is usually where teams need rule-based alerts, not just human inspection.

Currency is timeliness. A record can be correct and still be too old to use, especially in longitudinal work where status changes matter. A stale timestamp or delayed update deserves its own check, because stale data often masquerades as complete data.

The reading on data quality dimensions is useful if your team needs a shared vocabulary before writing rules. Teams often overweight correctness and undercheck plausibility and currency, which is exactly where the quiet failures hide.

Usefulness test: if a check can't tell you what kind of problem it found, it isn't sharp enough yet.

Attaching Validation to Every Handling Step

A five-step diagram showing a data validation workflow including transcription, transfer, update, save, and review processes.

Quality control belongs at every point where data are touched, not just in a final review. Clinical research guidance is explicit that validation should happen when data are transcribed, transferred, updated, or saved to a new medium, and the classic procedures include double data entry, programmatic range and consistency checks, regular error-rate review, and manager or peer review. The point is auditability, because each handling step is a place where error can be introduced or concealed.

A paper form moving into analysis

Start with paper forms. During transcription, one person enters the record and another rechecks a sample or performs double entry on fields that are easy to mistype, like dates or numeric measurements. A simple discrepancy log pays for itself fast here.

When the data move into a spreadsheet, run range checks and consistency rules immediately. A birth date in the future, a missing visit code, or a lab value outside the permitted domain should stop the file before anyone starts analysis.

When the spreadsheet becomes an analysis database, check the transfer itself. That means comparing row counts, key identifiers, and any fields that are vulnerable to truncation, recoding, or type conversion. If the database is the first place a problem appears, you've already lost the chance to tell whether the issue came from entry or migration.

After updates, route changes through peer or manager review. That review doesn't need to be ceremonial, it needs to answer one question, did the update preserve the meaning of the record?

The validation rules and continuous checks page is a useful example of how teams operationalize these guardrails in systems that run every day. The key is to tie the check to the step, because a periodic sweep that happens after three handoffs is already late.

A check that happens after the transfer is better than nothing, but it's not the same as controlling the transfer itself.

Measurable Thresholds for Quality Monitoring

Quality control becomes operational once the team agrees on numbers, not slogans. Defensible thresholds, named owners, and a fixed review cadence turn quality metrics into controls. A dashboard without a review owner is just wallpaper.

Sampling still matters. The quality-assurance guidelines support hard limits that teams can act on. At the site level, no more than 5% of enrolled participants should fail inclusion or exclusion criteria, enrollment should stay at least 90% of goal on time, the drop-out rate should stay at or below 5%, and the data-entry error rate should stay at or below 0.001%.

Metric

Threshold

How to Monitor

Random source recheck

5% of records

Reconcile selected records against source documents

Inclusion or exclusion failures

No more than 5% at a site

Review screening logs and protocol deviations

Enrollment progress

At least 90% of goal on time

Track enrollment against planned milestone dates

Drop-out rate

No greater than 5%

Monitor retention and withdrawal logs

Data-entry error rate

No greater than 0.001%

Compare entered values with source fields and log discrepancies

The point is not to collect these measures and admire them later. It is to review them on a fixed schedule, assign one owner for each threshold, and write down the response path before the first breach. That is where a timeliness monitoring view helps, because a delayed check is often the same as no check at all.

Keep the response proportional to the failure mode. A small run of entry errors calls for targeted rechecks, while repeated misses on enrollment or retention usually mean the process itself needs to change. If the threshold does not trigger a clear action, it is only decoration.

Provenance Documentation and Audit Trails

The question an auditor asks is never abstract. It's usually, “Where did this number come from?” If your answer requires memory, side conversations, and a hunt through old exports, the process isn't auditable yet. Provenance documentation is the record of origin, handling, and transformation that lets you answer that question without a scramble.

A defensible provenance record tells you who touched the data, what changed, when it changed, and why. It also preserves the original value when a correction is made, because the original is often the only way to understand whether a fix was valid or just convenient. That's especially important when multiple people have edited the same field across multiple tools.

What the audit trail needs to contain

The cleanest structure is simple. Keep version-controlled data files, a change log tied to each handling step, written queries for suspicious or missing values, and an explicit statement of unrecoverable missingness in the analysis plan. Those written queries matter because they turn uncertainty into a traceable decision instead of a buried exception.

The research guidance on iterative data quality control is blunt about this. Teams should run simple statistical reviews during collection, issue written queries for suspicious or missing values, and explicitly report unrecoverable missingness in the analysis plan (iterative workflow guidance). That's not bureaucratic overhead, it's how you keep the analysis honest about what can't be repaired.

If you can't show the path from source to analysis, you don't have a trail, you have a story.

A peer reviewer asking “where did this number come from” should be able to see the source record, the transformation applied, the date of the change, and the person who approved it. If the answer is a reconstruction project, the trail is too thin for serious work. Provenance is what lets the team explain the number instead of defending a memory of it.

Where AI Fits in Data Quality Control

AI is useful in data quality control, but only in a narrow lane. It can surface anomalies, cluster suspicious records, and spot patterns that are hard to encode as rules, which matters when the dataset is large and the baseline is stable. It also cuts down human triage time when it sits beside deterministic validation instead of replacing it. For teams building that layer, how AI detects data anomalies in data pipelines is the right model to study.

The risk is that automated cleaning can invent plausible-looking corrections that are wrong. Recent academic coverage points to a real gap in standardized benchmarks for LLM-based data quality systems, and it warns that automated cleansing can introduce hallucinated corrections. Treat AI as a reviewer's assistant, not as the authority that rewrites records on its own.

Safe uses and unsafe uses

Safe uses include flagging anomalies for review, prioritizing records for human inspection, and suggesting candidate issues that a person can confirm. Unsafe uses include overwriting source values without disclosure, resolving ambiguous records without human sign-off, and acting as the sole gatekeeper for release.

The guardrail is boring but effective. Keep original values alongside any automated corrections, sample-review the AI-flagged records, and document the decision path for every accepted change. Store the original value in a separate column with the transformation timestamp and the approver's name so every correction stays reversible and auditable. If your team wants a broader governance view, the governance and AI data tools resource from MakeAutomation is a useful complement because it frames AI inside operational controls rather than hype.

For teams that want to apply AI without surrendering control, a platform like digna can sit in the validation and anomaly layer, where checks stay inspectable and tied to the underlying data. That is the right pattern. AI helps you notice, humans decide.

Your Data Quality Control Checklist

A checklist of five essential data quality control steps for researchers to verify and maintain data accuracy.

A checklist only works if it is short enough to use during live work and strict enough to catch drift before analysis starts. Keep it on the bench, not in a slide deck.

  • Random source recheck: Recheck a random 5% sample against source material on a fixed weekly schedule, so transcription drift gets caught while the file is still open.

  • Validation at every handoff: Attach a specific control to transcription, transfer, update, save, and review. A general promise to “check later” does not survive a busy pipeline.

  • Threshold owner named: Assign one person to review enrollment, dropout, and entry-error thresholds on schedule. Metrics change behavior only when someone is clearly accountable for them.

  • Provenance recorded: Keep version-controlled files, a change log, and written queries for suspicious or missing values tied to the exact transformation. That makes later audits possible without reconstructing the whole chain from memory.

  • AI corrections reviewed: Require human review of AI-flagged records, keep the original value, and log why a correction was accepted or rejected.

  • Cross-database consistency: Compare key fields across systems so concordance problems surface before publication, not after.

If you want a practical operational model, pair this checklist with a live control layer instead of another spreadsheet. A governance and AI data tools setup from governance and AI data tools can keep record-level validation, anomaly detection, timeliness monitoring, and schema tracking inside the environment where the data already lives. digna fits that kind of control loop when research teams need checks that stay inspectable and tied to the underlying records.

✦ Généré avec l'intelligence artificielle

Partager sur X
Partager sur X
Partager sur Facebook
Partager sur Facebook
Partager sur LinkedIn
Partager sur LinkedIn

Rencontrez l'équipe derrière la plateforme

Une équipe viennoise d'experts en IA, en données et en logiciel, portée

par la rigueur académique et l'expérience de l'entreprise.

Rencontrez l'équipe derrière la plateforme

Une équipe viennoise d'experts en IA, en données et en logiciel, portée par la rigueur académique et l'expérience de l'entreprise.

Produit

Intégrations

Ressources

Société

INDEXED BYIndexerNow INDEXED BYIndexerNow