Chest radiograph screening · open to evaluation
Chest X-ray screening,
built to be checked.
Four independently trained models flag pneumonia, mark the region they responded to, and explain the result in writing. Every figure we publish was measured on scans the models had never seen — so a department can judge it and a person can understand it.
Datasets and published models this is built from
- Stanford CheXpert
- MIMIC-CXR · MIT & Harvard
- PadChest · Alicante
- RSNA Pneumonia Challenge
- IEEE TMI segmentation
Us and everyone else
Six questions. We answer all six.
We put six questions to the three most visible consumer chest X-ray tools we could find, on 31 August 2026, and to ourselves. We are not naming them — what a marketing page says today it may not say next month, and we are not going to make a claim about someone else that we cannot keep true. Take the questions and ask them of anyone, including us.
The three we checked
0 / 6
questions answered in public
PneumonAI
0 / 6
answered with a figure and the sample it came from
The three we checked
PneumonAI
What is its accuracy, and on which images?
The three we checked
No figure
None of the three published a single accuracy figure.
PneumonAI
Two figures, two sample sizes
AUC 0.81 on 26 335 scans with radiologist-drawn labels, 0.72 on 1 082 with labels parsed from reports — both with sample sizes, both on the research page.
Which models and datasets is it built on?
The three we checked
Not named
Not named. One says only that it uses a general-purpose language model.
PneumonAI
Named in full
Every model, dataset, paper and licence named, including the one we trained ourselves and why we had to.
What does the number actually mean?
The three we checked
Unexplained
Not explained anywhere we could find.
PneumonAI
Explained three times over
It is a calibrated agreement score, not a probability of disease — explained on the homepage, in the FAQ, and on the model page.
When is it wrong, and how often?
The three we checked
No number
"Not infallible", and nothing more specific.
PneumonAI
Failure rates published
It marks no region on 18% of positive scans, declines about 1 scan in 35, and flags roughly half of the scans that turn out to be clear.
Has a regulator cleared it?
The three we checked
Not stated
Not stated.
PneumonAI
Answered, and the answer is no
No. No CE mark, no FDA clearance, no clinical validation — said on the first screen you reach as a clinic.
What happens to the image I upload?
The three we checked
Unclear
Two say nothing. One stores it in an account; whether it trains on it is not stated.
PneumonAI
Deleted in ten minutes
Deleted about ten minutes after the result, never used for training, and a DICOM header never leaves your device at all.
One of the three claims its AI surpasses human radiologists. It publishes no number behind that sentence. We would rather be the site with the lower figure and the arithmetic attached.
See it run
One scan, start to finish
Recorded in one take on this site — a browser driving itself, not a screen capture. The scan is from NIH ChestX-ray14, a public de-identified research set, not from a patient of ours. The wait for the written report is sped up 6× and labelled where it happens; nothing else is edited.
What is inside
What each one is measured on
Pneumonia ensemble
3 models, 3 datasets, 3 institutions
DetailsWhereFinding localizer
Peak lands in a radiologist's box 86% of the time — 8.1× chance
DetailsWhat is lungLung segmentation
Traces both lungs at 0.97 Dice — published in IEEE TMI
DetailsIs this a chest X-rayX-ray validator
Checked on 25 images — a spot check, not a test set
DetailsLive today
One condition, measured end to end and published with its numbers.
In development
We are extending the ensemble one condition at a time. A condition goes live only once we can measure it the way pneumonia is measured on this page — on scans the models have never seen, with the result published whether or not it flatters us.
Numbers you can check
Everything here was measured, and we show you on what
On 26 335 chest X-rays labelled by radiologists. Members: PadChest 0.78, Stanford 0.77, MIMIC 0.70.
5 850 annotated positives, at the operating point τ = 0.45.
Which means 43 of every 100 clear scans do get flagged. For screening a missed case costs more than a false alarm, so this is the trade — but if you are flagged, the likeliest explanation is still that you do not have pneumonia.
Against 10.6% for a point placed at random.
And what it does not do. The localizer marks no region at all on 18% of scans that do carry a finding — the verdict stands, but there is nothing to point at. Roughly one scan in thirty-five is declined outright because the segmentation could not find lung fields in it. Neither is a bug; both are rates you will meet in an evaluation, so they are here rather than a click away.
How to read the score
Each of the three models starts flagging a scan at a different point — one of them as low as 0.044 — so their raw outputs cannot be compared, let alone averaged. Every score is rescaled to put that model’s own flagging point at 0.5, and the average is measured against a line at 0.45. Read it as reaction strength: how far past their own line the models landed. The ensemble’s line sits below 0.5 on purpose: it is tuned to flag slightly more readily than its members would on average, because for screening a missed case costs more than a false alarm.
Why you will see two numbers
The same ensemble, on the same images, scores 0.81 against regions drawn by radiologists and 0.72 against NIH ChestX-ray14, labels parsed from report text. Nothing about the models changes between those runs — only who wrote the answer key. We publish both, because a gap you can explain is worth more than one round number you cannot.
For clinics and departments
Everything you need to judge it before you trust it
PneumonAI is open for evaluation, not deployment. There is no clearance behind it, and the case we make is not that you should believe the result — it is that you can check it, on your own scans, against your own reports, with every measurement we hold.
Run your own cases through it
Nothing to install, no account, no procurement. Put your own scans in and compare the verdict and the marked region against the report you already have for them.
Every number, with the conditions attached
Sample sizes, datasets, operating points, and the runs where our own figure came out worse than we hoped — including the measurement we corrected downwards in public after finding our test set contaminated.
Your images do not become our dataset
DICOM is opened in the browser and only the pixels reach us — the header, with the patient name and ID in it, never leaves your device. The image is deleted about ten minutes after the result, a written report lives seven days, nothing is used for training, and there are no cookies and no analytics.
What it is not
Not a medical device. No CE mark, no FDA clearance, no clinical validation. It must not sit in a diagnostic pathway, and any evaluation should be read as research rather than as a second opinion.
Where this is going
An API and a workflow integration are what we are building toward, so a result can reach the place a radiologist already works instead of a browser tab. Neither exists yet, and we would rather say so than publish a page describing one. If that is the part you care about, tell us — it decides what gets built first.
Before you upload
What the result means, what happens to your scan, and how the report is opened
Six questions people ask first, answered in full — including the one that matters most, which is how to read the number you get back.
