Data Verification
Independent verification of every received dataset (completeness, label fidelity, de-identification, subgroup representativeness) before a single training epoch runs.
Overview
Training on a defective dataset produces a defective model and a year of confused debugging. Most defects are detectable on receipt, if anyone is actually looking. Most teams discover them six months later, when validation fails.
We run independent verification on every received dataset against the original specification. Defects are documented, reported back to the vendor under contractual remediation, and your training team only sees data that passed.
Our Process
-
1
Acceptance criteria recap
Pull the criteria from the dataset specification and the data use agreement.
-
2
Completeness verification
Volume, modality coverage, DICOM tag completeness, label coverage.
-
3
Label fidelity sampling
Stratified re-read of a sample by independent annotators.
-
4
De-identification audit
Automated scan plus manual review of edge cases.
-
5
Representativeness check
Distribution across demographics and clinical subgroups vs. specification targets.
-
6
Acceptance or rejection
Defects logged with remediation request; accepted data flows to training pipeline.
Frequently Asked Questions
Who reads labels for fidelity sampling?
Board-certified specialists in the relevant modality, recruited through a managed network.
Does this overlap with your QMS?
Verification findings feed your QMS records, Design Verification per IEC 62304 and data quality per ISO 13485.
Can this catch dataset drift?
Yes, verification on each new delivery surfaces drift against earlier deliveries.
What if labels are subjective?
We document inter-rater agreement targets and accept against the specified threshold.
Train on data you have actually inspected.
Send us your most recent dataset delivery and acceptance criteria. We will return a verification report within ten business days.
Start a Conversation