What the programme covers
Training data must contain enough relevant variation for the intended use without leaking repeated patients, derived copies or validation cases across splits. A MedCorpora programme makes scanner, protocol, site, population and label distributions explicit before model use.
What a useful specification includes
A defensible request defines the clinical task, source evidence and acceptance criteria before patient-level data moves. The exact fields and thresholds depend on the intended model claim.
- Prediction, detection, segmentation or representation-learning objective
- Case definition and inclusion or exclusion logic
- Class balance and difficult-case strategy
- Patient-level train, tune and test separation
- Label source, reviewer qualifications and uncertainty policy
- Required scanner, protocol, site and population coverage
Quality and validation controls
The release manifest records how each case was counted, which transformations ran and why cases were excluded. Training labels remain connected to their source evidence and review state.
- Duplicate, near-duplicate and repeat-examination detection
- Missing sequence, view or clinical-field analysis
- Class and subgroup distribution reporting
- Label noise, disagreement and unknown-state handling
- Versioned preprocessing and dataset lineage
Availability, rights and delivery
Training cohorts can be released in stages so the buyer can test loaders, labels and technical assumptions before final-scale processing.
Public pages describe a sourcing and engineering capability, not guaranteed ready inventory. Each release remains subject to verified programme inventory, programme-specific authorization, privacy review, technical acceptance and buyer licence terms.
Questions, answered directly.
Can training data include negative cases?+
Yes. Negative and normal definitions must be explicit and supported by an appropriate reference standard.
Can MedCorpora help create dataset splits?+
Yes. Splits can be materialized at patient level with site, scanner, time or population constraints.
Does more data always improve a model?+
No. Relevance, label quality, variation and leakage control can matter more than unqualified volume.
Can the data train commercial models?+
Only when the programme authority and buyer licence expressly cover the intended commercial use.