All models

Trained in-house

Finding localizer

The only model here we trained ourselves. It answers one question — where on the image the most notable region is — and deliberately has no vote on whether the scan is flagged.

Measured

86%
Peak inside a radiologist's box
Against 11% for a point placed at random — 8.1× chance.
8.1×
Better than chance
The three classifiers' own heatmaps managed 1.3-2.7× on the same scans, which is why locating a finding is this model's job and not theirs.
18%
Positives with no region marked
Not a failure and not a missing measurement. The result says so rather than showing an empty picture, and it does not soften the verdict.

Measured on held-out scans the model never saw during training. The first run reported 92.4%, and that figure was memory rather than quality — 204 of the 250 test positives were in its own training set. The split is now made by patient with the same seed the training notebook uses, and both scripts read one shared list so the comparison cannot drift.

Why we trained it

Models that do this already exist — ChEX, BioViL, Foundation X — and every one of them was unusable here for the same reason, which is the licence rather than the quality. BioViL is published under MIT and its own model card still says that any deployed use, commercial or otherwise, is out of scope. ChEX ships weights but was trained on MIMIC-CXR, MS-CXR and VinDr-CXR, all of which are PhysioNet credentialed data, research and education only.

The RSNA Pneumonia Detection Challenge is the one chest X-ray dataset with drawn regions that permits commercial use with attribution. That makes the chain clean end to end: our code, permitted data, our weights. The encoder starts from ImageNet rather than from our own PadChest classifier for the same reason — on held-out data the two were indistinguishable, so the tie was broken on provenance.

What the map is, and is not

The highlighted area is where the model responded most strongly. Its edges are that response falling off, not a measured boundary of anything, and its shape says nothing about the shape of what is there. It runs at 512×512 rather than the classifiers' 224 precisely so that the map is not a coarse grid stretched over a chest — the resolution is the feature.

Technical details

Architecture
U-Net (DenseNet121 encoder, ImageNet init)
Library
segmentation-models-pytorch
Input
1 × 1 × 512 × 512
Training data
RSNA Pneumonia Detection Challenge (2018)
Size
26 684 images, 9 555 radiologist-drawn boxes
Trained by
PneumonAI, 2026-08-22
License
MIT (code); dataset permits commercial use with attribution

Reference and attribution

Radiological Society of North America. RSNA Pneumonia Detection Challenge (2018). Annotations by the RSNA and the Society of Thoracic Radiology.