Evidence

Every number here, and how we got it

Each figure below was measured on our own pipeline, on images the models had not seen, and is stated with the conditions behind it. Where a measurement is still open, we say so rather than filling the gap with a rounder number.

How the ensemble was measured

0.81
Ensemble ROC AUC
On radiologist-drawn labels. Members alone: PadChest 0.78, Stanford 0.77, MIMIC 0.70.
90%
Sensitivity
At τ = 0.45, the operating point in production.
57%
Specificity
The rest of the clear scans are flagged. Deliberate: for screening, a missed case costs more than a false alarm.

Measured on 26 335 chest X-rays from the RSNA Pneumonia Detection Challenge, regions drawn by radiologists — 5 850 of them annotated positive — run through the production pipeline rather than a notebook reimplementation of it.

The threshold was chosen on a measurement run and not before it. τ = 0.45 rather than the 0.50 that maximises Youden’s J: that statistic weighs a missed case and a false alarm equally, which for screening is the wrong trade.

Two answer keys, one ensemble

LabelsScansAUCSensitivitySpecificity
Drawn by radiologists (RSNA)26 3350.8190%57%
Parsed from report text (NIH)1 0820.7285%46%

Same models, same weights, same threshold, and largely the same images — RSNA is a re-annotation of the NIH collection. The only thing that differs between those two rows is who wrote the answer key. NIH’s pneumonia labels were extracted from the text of radiology reports by a parser, and their noisiness is a documented property of that dataset; RSNA had radiologists mark the regions by hand for this exact task.

We report the expert-label figure as the headline because it measures the question the product is actually asked, and the parser-label figure alongside it because a number without its answer key is not a measurement. Neither run is on data any of the three classifiers were trained on.

How the localizer was measured

86%
Peak inside a radiologist's region
Against 11% at random — 8.1× chance.
1.3–2.7×
What the classifiers' own heatmaps managed
On the same scans. This gap is why locating a finding is a separate model's job.
18%
Positives with nothing marked
Reported as such rather than shown as an empty picture.

The first run of this measurement reported 92.4%, and that number was memory rather than quality: 204 of the 250 test positives were in the model's own training set. The split is now made by patient with the same seed the training notebook uses, and the two scripts read one shared patient list so the comparison cannot drift apart. We record the wrong number here because a measurement harness that has never been checked against a known answer is as good a source of false confidence as a model.

Open questions we are still working on

The localizer was measured on RSNA data only. How it behaves on images from equipment it has never seen is unknown, and no second box-annotated chest X-ray dataset of comparable size exists to check it against. This is the largest open question about the product and it is not resolvable with the data that exists publicly.

Positive predictive value at real-world prevalence has not been re-measured for the current ensemble; the figure in our earlier report describes the combination that preceded it. The balanced sample used for the threshold is adequate for choosing one — 539 positives, ±2.5 points on sensitivity — but is not the full split.

The validator does not separate a frontal view from a lateral one, or an adult chest from a child's. Both pass, and both are outside the distribution the classifiers were trained on.

Datasets

DatasetInstitutionYearUsed for
PadChestSan Juan Hospital, Alicante2020Ensemble member
Stanford CheXpertStanford ML Group2019Ensemble member
MIMIC-CXRMIT + Harvard BIDMC2019Ensemble member
RSNA Pneumonia DetectionRadiological Society of North America2018Localizer training (ours)
NIH ChestX-ray14US National Institutes of Health2017Held-out measurement only

Why an ensemble

Each classifier was trained at a different institution, on different patients, by a different team. That is the entire argument: models that fail in the same places would agree with each other and tell you nothing, so agreement between these three carries information that agreement between three copies of one model would not.

Averaging their raw outputs would be meaningless — their flagging points differ by more than an order of magnitude. Each score is first rescaled so that 0.5 falls exactly on that model's own flagging point, and only then averaged.

Licensing

The localizer is the one model here we trained ourselves, and the reason is the licence rather than the quality. Published localization models exist — ChEX, BioViL, Foundation X — and each is unusable in a deployed product: BioViL is released under MIT while its own model card places any deployed use out of scope, and ChEX was trained on PhysioNet credentialed data marked research and education only. The RSNA challenge is the one box-annotated chest X-ray dataset that permits commercial use with attribution.

Radiological Society of North America. RSNA Pneumonia Detection Challenge (2018). Annotations by the RSNA and the Society of Thoracic Radiology.

The same questions, asked of anyone

Everything above is our answer to six questions worth putting to any tool that reads a chest X-ray. We put them to the tools we could find as well, and published what came back — without naming them, because a marketing page can change and we are not going to make a claim about someone else that we cannot keep true.

References