Evidence
Every number here, and how we got it
Each figure below was measured on our own pipeline, on images the models had not seen, and is stated with the conditions behind it. Where a measurement is still open, we say so rather than filling the gap with a rounder number.
How the ensemble was measured
Measured on 26 335 chest X-rays from the RSNA Pneumonia Detection Challenge, regions drawn by radiologists — 5 850 of them annotated positive — run through the production pipeline rather than a notebook reimplementation of it.
The threshold was chosen on a measurement run and not before it. τ = 0.45 rather than the 0.50 that maximises Youden’s J: that statistic weighs a missed case and a false alarm equally, which for screening is the wrong trade.
Two answer keys, one ensemble
| Labels | Scans | AUC | Sensitivity | Specificity |
|---|---|---|---|---|
| Drawn by radiologists (RSNA) | 26 335 | 0.81 | 90% | 57% |
| Parsed from report text (NIH) | 1 082 | 0.72 | 85% | 46% |
Same models, same weights, same threshold, and largely the same images — RSNA is a re-annotation of the NIH collection. The only thing that differs between those two rows is who wrote the answer key. NIH’s pneumonia labels were extracted from the text of radiology reports by a parser, and their noisiness is a documented property of that dataset; RSNA had radiologists mark the regions by hand for this exact task.
We report the expert-label figure as the headline because it measures the question the product is actually asked, and the parser-label figure alongside it because a number without its answer key is not a measurement. Neither run is on data any of the three classifiers were trained on.
How the localizer was measured
The first run of this measurement reported 92.4%, and that number was memory rather than quality: 204 of the 250 test positives were in the model's own training set. The split is now made by patient with the same seed the training notebook uses, and the two scripts read one shared patient list so the comparison cannot drift apart. We record the wrong number here because a measurement harness that has never been checked against a known answer is as good a source of false confidence as a model.
Open questions we are still working on
The localizer was measured on RSNA data only. How it behaves on images from equipment it has never seen is unknown, and no second box-annotated chest X-ray dataset of comparable size exists to check it against. This is the largest open question about the product and it is not resolvable with the data that exists publicly.
Positive predictive value at real-world prevalence has not been re-measured for the current ensemble; the figure in our earlier report describes the combination that preceded it. The balanced sample used for the threshold is adequate for choosing one — 539 positives, ±2.5 points on sensitivity — but is not the full split.
The validator does not separate a frontal view from a lateral one, or an adult chest from a child's. Both pass, and both are outside the distribution the classifiers were trained on.
Datasets
| Dataset | Institution | Year | Used for |
|---|---|---|---|
| PadChest | San Juan Hospital, Alicante | 2020 | Ensemble member |
| Stanford CheXpert | Stanford ML Group | 2019 | Ensemble member |
| MIMIC-CXR | MIT + Harvard BIDMC | 2019 | Ensemble member |
| RSNA Pneumonia Detection | Radiological Society of North America | 2018 | Localizer training (ours) |
| NIH ChestX-ray14 | US National Institutes of Health | 2017 | Held-out measurement only |
Why an ensemble
Each classifier was trained at a different institution, on different patients, by a different team. That is the entire argument: models that fail in the same places would agree with each other and tell you nothing, so agreement between these three carries information that agreement between three copies of one model would not.
Averaging their raw outputs would be meaningless — their flagging points differ by more than an order of magnitude. Each score is first rescaled so that 0.5 falls exactly on that model's own flagging point, and only then averaged.
Licensing
The localizer is the one model here we trained ourselves, and the reason is the licence rather than the quality. Published localization models exist — ChEX, BioViL, Foundation X — and each is unusable in a deployed product: BioViL is released under MIT while its own model card places any deployed use out of scope, and ChEX was trained on PhysioNet credentialed data marked research and education only. The RSNA challenge is the one box-annotated chest X-ray dataset that permits commercial use with attribution.
Radiological Society of North America. RSNA Pneumonia Detection Challenge (2018). Annotations by the RSNA and the Society of Thoracic Radiology.
The same questions, asked of anyone
Everything above is our answer to six questions worth putting to any tool that reads a chest X-ray. We put them to the tools we could find as well, and published what came back — without naming them, because a marketing page can change and we are not going to make a claim about someone else that we cannot keep true.
References
- Bustos et al., Medical Image Analysis 2020 — "PadChest: A large chest x-ray image dataset with multi-label annotated reports"
- Irvin et al., AAAI 2019 — "CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels"
- Johnson et al., Nature Scientific Data 2019 — "MIMIC-CXR: A large publicly available database of labeled chest radiographs"
- Cohen et al., PMLR 2022 — "TorchXRayVision: A library of chest X-ray datasets and models"
- Lian et al., IEEE TMI 2021 — "A Structure-Aware Relation Network for Thoracic Diseases Detection and Segmentation" (lung segmentation, Dice 0.97)
- Shih et al., Radiology: Artificial Intelligence 2019 — Augmenting the NIH Chest Radiograph Dataset with Expert Annotations of Possible Pneumonia (the annotations our localizer was trained on)
