Representative Clinical Datasets

GMLP Principle 3. Make datasets representative of the patients your device will actually be used on: across age, sex, race, ethnicity, geography, and condition.

Overview

Unrepresentative datasets produce devices that work for some patients and silently underperform for others. Regulators are increasingly explicit about subgroup performance, and rightfully so.

We design and audit datasets for representativeness from sourcing through validation, with subgroup performance evaluated as a first-class outcome.

Our Process

  1. 1

    Subgroup target setting

    Per-subgroup target volumes derived from intended use.

  2. 2

    Sourcing strategy

    Data partners selected for subgroup coverage.

  3. 3

    Continuous monitoring

    Track representativeness across data ingest.

  4. 4

    Subgroup performance evaluation

    Per-subgroup performance metrics in validation.

  5. 5

    Drift management

    Detect and respond to representativeness drift.

Frequently Asked Questions

How do we set targets without overweighting rare subgroups?

Power analysis grounded in clinical impact and submission requirements.

What if data is genuinely scarce for some subgroups?

Documented limitation, mitigations in labeling, and post-market monitoring focus.

Does this delay our timeline?

Sometimes. But late discovery is more expensive than early planning.

FDA expectations on subgroup performance?

Increasingly explicit. We align with current draft and final guidance.

Make datasets representative on purpose.

Send us your intended use and current datasets. We will return a representativeness assessment within four weeks.

Start a Conversation