Hong Kong Healthcare Artificial Intelligence SocietyHong Kong Healthcare Artificial Intelligence Society

FDA AI Device TPLC: Data, Models & Validation Terminology

FDA's draft expectations for data management, model description and development, and the critical terminology differences between FDA 'validation' and AI community usage.

Clinical data pipelines and AI model development for medical device validation

For AI-enabled devices, the model is part of the mechanism of action. FDA's draft guidance therefore devotes substantial attention to data management, model description and development, and clear validation terminology. Hong Kong clinicians evaluating vendor evidence should understand these distinctions — marketing materials often use AI research language that does not align with regulatory definitions.

Data management

Data management practices — how data are collected, processed, annotated, stored, controlled, and used — are critical because AI performance depends heavily on data quality, diversity, and quantity.

Key submission themes

Data collection should describe sites, time periods, inclusion/exclusion criteria, real-world data (RWD) fit-for-purpose assessments, quality assurance, dataset size, diversity enrollment mechanisms, and any synthetic data with justification.

Data cleaning and processing must be documented for development data. Test data should only be processed in ways representative of real-world intended use, aligned with the final AI-DSF pre-processing.

Reference standards (ground truth) should reflect the clinical task, with documented establishment methods, uncertainty, equivocal case handling, and clinician grading protocols including blinding and inter-/intra-observer variability.

Data annotation should describe annotator expertise, instructions, quality/consistency evaluation, and correction plans.

Independence between development and test data is essential. Test data should generally come from sites different from training data, be sequestered from developers, and support robust external validation. Data leakage between sets inflates reported performance.

Representativeness requires explaining how data reflect the intended use population — disease spectrum, demographics (sex, age, race, ethnicity), acquisition equipment, and clinical settings. Subgroup analyses and covariate distributions should be provided. When outside-U.S. (OUS) data are used, sponsors should explain comparability to the target population and medical practice.

Bias and confounding

FDA emphasises that AI bias — systematic but sometimes unforeseeable incorrect results — can arise from non-representative training data, confounders (e.g., all diseased cases imaged on one scanner), or under-represented subgroups. The same confounders in both development and validation data can hide spurious correlations.

For Hong Kong settings, ask whether validation included Asian or local populations, diverse acquisition equipment, and multiple sites — not only U.S.-centric datasets.

Model description and development

The software description should enable a competent AI practitioner to understand each model:

  • Inputs, outputs, architecture, features, feature selection, loss functions, parameters
  • Customisable technical elements and quality control for input data
  • Pre-processing, post-processing, augmentation, or synthesis methods

Model development documentation should cover training methods, paradigms (supervised, federated, etc.), regularisation, hyperparameters, convergence curves, tuning evaluation, pre-trained models, ensemble methods, threshold/operating point selection, and output calibration.

When multiple models combine, diagrams showing how outputs merge into device outputs are encouraged.

For models that update after deployment, FDA encourages early Q-Submission engagement and review of PCCP guidance.

Validation terminology: FDA vs AI community

This guidance explicitly resolves terminology conflicts:

TermFDA / device contextCommon AI research usage
ValidationConfirmation with objective evidence that intended-use requirements are consistently fulfilled (21 CFR 820.3(z))Sometimes used for tuning or hold-out testing during development
DevelopmentTraining, tuning, and tuning evaluation ("internal testing")Same general concept
Test dataData for verification and validation — not part of developmentOften overlaps with "validation set" in ML workflows

Practical rule for submissions: Do not use "validation" to describe training or tuning. Performance validation evaluates the model on independent datasets under clinically relevant conditions.

Model version history

Submissions should document which model version was tested, differences from the released version, and safeguards against post-hoc adjustment to test data without regulatory concurrence. New UDIs may be required for new model versions where UDI rules apply.

Source: U.S. FDA — Artificial Intelligence-Enabled Device Software Functions: Lifecycle Management and Marketing Submission Recommendations (January 2025, draft guidance)

Ready to test your knowledge?

Take a short quiz based on this article to check your understanding.

Take the quiz