Representative Clinical Datasets
GMLP Principle 3. Make datasets representative of the patients your device will actually be used on: across age, sex, race, ethnicity, geography, and condition.
Overview
Unrepresentative datasets produce devices that work for some patients and silently underperform for others. Regulators are increasingly explicit about subgroup performance, and rightfully so.
We design and audit datasets for representativeness from sourcing through validation, with subgroup performance evaluated as a first-class outcome.
Our Process
-
1
Subgroup target setting
Per-subgroup target volumes derived from intended use.
-
2
Sourcing strategy
Data partners selected for subgroup coverage.
-
3
Continuous monitoring
Track representativeness across data ingest.
-
4
Subgroup performance evaluation
Per-subgroup performance metrics in validation.
-
5
Drift management
Detect and respond to representativeness drift.
Frequently Asked Questions
How do we set targets without overweighting rare subgroups?
Power analysis grounded in clinical impact and submission requirements.
What if data is genuinely scarce for some subgroups?
Documented limitation, mitigations in labeling, and post-market monitoring focus.
Does this delay our timeline?
Sometimes. But late discovery is more expensive than early planning.
FDA expectations on subgroup performance?
Increasingly explicit. We align with current draft and final guidance.
Make datasets representative on purpose.
Send us your intended use and current datasets. We will return a representativeness assessment within four weeks.
Start a Conversation