All models

Ensemble · 3 models

Pneumonia Detection

Three DenseNet121 classifiers, each trained on a different hospital dataset — Stanford CheXpert, MIMIC-CXR, PadChest — read the scan independently and vote. Alone they reach 0.78, 0.77 and 0.70 AUC on our test set; together, 0.81.

Why an ensemble

Three models trained on data from three different institutions, on three different continents, by three different teams. That is the whole argument: models that fail in the same places would agree with each other and tell you nothing, so agreement between these three carries information that agreement between three copies of one model would not.

Whether the spread between the three predicts anything about correctness is something we have not measured. Until we have, read a split vote as a split vote — on our test run all three agreed to flag about half the time, so disagreement is the normal case rather than a warning sign.

How the score is built

Each model starts flagging a scan at a different raw score — 0.0444, 0.0775, 0.0229. Averaging those directly would be meaningless, so each is first rescaled to put that model’s own flagging point at exactly 0.5. The three rescaled scores are averaged and compared against a fixed line at 0.45.

That line sits below 0.5, and deliberately: the ensemble is tuned to flag slightly more readily than its members would on average, because for screening a missed case costs more than a false alarm. It was chosen on a measurement run rather than picked, and not at the value that maximises Youden’s J — that statistic weighs a missed case and a false alarm equally, which for screening is the wrong trade.

The result is not a probability of pneumonia and cannot be read as one. It is how far past their own flagging points the three models landed, on average. Every model takes a 1 × 1 × 224 × 224 grayscale input.

The three models

Model 1 · Alicante
Dataset
PadChest
Institution
Hospital Universitario de San Juan, Alicante (Spain)
Architecture
DenseNet121
Input
1 × 1 × 224 × 224
Op. threshold (raw)
0.0444
License
Apache 2.0
Model 2 · Stanford
Dataset
Stanford CheXpert
Institution
Stanford ML Group
Architecture
DenseNet121
Input
1 × 1 × 224 × 224
Op. threshold (raw)
0.0775
License
Apache 2.0
Model 3 · MIT / Harvard
Dataset
MIMIC-CXR
Institution
MIT + Harvard (Beth Israel Deaconess Medical Center)
Architecture
DenseNet121
Input
1 × 1 × 224 × 224
Op. threshold (raw)
0.0229
License
PhysioNet Credentialed Health Data License 1.5.0

Built with

All three models are served through the torchxrayvision library.