Chest radiograph screening · open to evaluation

Chest X-ray screening,
built to be checked.

Four independently trained models flag pneumonia, mark the region they responded to, and explain the result in writing. Every figure we publish was measured on scans the models had never seen — so a department can judge it and a person can understand it.

Evaluating it for a clinic?See how it worksFree while we gather feedback · no account

Upload your X-ray

PNG, JPEG or DICOM · up to 10 MB

Datasets and published models this is built from

  • Stanford CheXpert
  • MIMIC-CXR · MIT & Harvard
  • PadChest · Alicante
  • RSNA Pneumonia Challenge
  • IEEE TMI segmentation

Us and everyone else

Six questions. We answer all six.

Our answers in full

We put six questions to the three most visible consumer chest X-ray tools we could find, on 31 August 2026, and to ourselves. We are not naming them — what a marketing page says today it may not say next month, and we are not going to make a claim about someone else that we cannot keep true. Take the questions and ask them of anyone, including us.

The three we checked

0 / 6

questions answered in public

PneumonAI

0 / 6

answered with a figure and the sample it came from

What is its accuracy, and on which images?

The three we checked

No figure

None of the three published a single accuracy figure.

PneumonAI

Two figures, two sample sizes

AUC 0.81 on 26 335 scans with radiologist-drawn labels, 0.72 on 1 082 with labels parsed from reports — both with sample sizes, both on the research page.

Which models and datasets is it built on?

The three we checked

Not named

Not named. One says only that it uses a general-purpose language model.

PneumonAI

Named in full

Every model, dataset, paper and licence named, including the one we trained ourselves and why we had to.

What does the number actually mean?

The three we checked

Unexplained

Not explained anywhere we could find.

PneumonAI

Explained three times over

It is a calibrated agreement score, not a probability of disease — explained on the homepage, in the FAQ, and on the model page.

When is it wrong, and how often?

The three we checked

No number

"Not infallible", and nothing more specific.

PneumonAI

Failure rates published

It marks no region on 18% of positive scans, declines about 1 scan in 35, and flags roughly half of the scans that turn out to be clear.

Has a regulator cleared it?

The three we checked

Not stated

Not stated.

PneumonAI

Answered, and the answer is no

No. No CE mark, no FDA clearance, no clinical validation — said on the first screen you reach as a clinic.

What happens to the image I upload?

The three we checked

Unclear

Two say nothing. One stores it in an account; whether it trains on it is not stated.

PneumonAI

Deleted in ten minutes

Deleted about ten minutes after the result, never used for training, and a DICOM header never leaves your device at all.

One of the three claims its AI surpasses human radiologists. It publishes no number behind that sentence. We would rather be the site with the lower figure and the arithmetic attached.

See it run

One scan, start to finish

Try it on your own scan
pneumonai.org

Recorded in one take on this site — a browser driving itself, not a screen capture. The scan is from NIH ChestX-ray14, a public de-identified research set, not from a patient of ours. The wait for the written report is sped up 6× and labelled where it happens; nothing else is edited.

What is inside

What each one is measured on

All models

Live today

Pneumonia

One condition, measured end to end and published with its numbers.

In development

We are extending the ensemble one condition at a time. A condition goes live only once we can measure it the way pneumonia is measured on this page — on scans the models have never seen, with the result published whether or not it flatters us.

Numbers you can check

Everything here was measured, and we show you on what

0.81
Ensemble ROC AUC

On 26 335 chest X-rays labelled by radiologists. Members: PadChest 0.78, Stanford 0.77, MIMIC 0.70.

90%
Pneumonia cases caught

5 850 annotated positives, at the operating point τ = 0.45.

57%
Clear scans left unflagged

Which means 43 of every 100 clear scans do get flagged. For screening a missed case costs more than a false alarm, so this is the trade — but if you are flagged, the likeliest explanation is still that you do not have pneumonia.

86%
Localizer peak inside a marked region

Against 10.6% for a point placed at random.

And what it does not do. The localizer marks no region at all on 18% of scans that do carry a finding — the verdict stands, but there is nothing to point at. Roughly one scan in thirty-five is declined outright because the segmentation could not find lung fields in it. Neither is a bug; both are rates you will meet in an evaluation, so they are here rather than a click away.

How to read the score

Each of the three models starts flagging a scan at a different point — one of them as low as 0.044 — so their raw outputs cannot be compared, let alone averaged. Every score is rescaled to put that model’s own flagging point at 0.5, and the average is measured against a line at 0.45. Read it as reaction strength: how far past their own line the models landed. The ensemble’s line sits below 0.5 on purpose: it is tuned to flag slightly more readily than its members would on average, because for screening a missed case costs more than a false alarm.

Why you will see two numbers

The same ensemble, on the same images, scores 0.81 against regions drawn by radiologists and 0.72 against NIH ChestX-ray14, labels parsed from report text. Nothing about the models changes between those runs — only who wrote the answer key. We publish both, because a gap you can explain is worth more than one round number you cannot.

For clinics and departments

Everything you need to judge it before you trust it

PneumonAI is open for evaluation, not deployment. There is no clearance behind it, and the case we make is not that you should believe the result — it is that you can check it, on your own scans, against your own reports, with every measurement we hold.

Run your own cases through it

Nothing to install, no account, no procurement. Put your own scans in and compare the verdict and the marked region against the report you already have for them.

Every number, with the conditions attached

Sample sizes, datasets, operating points, and the runs where our own figure came out worse than we hoped — including the measurement we corrected downwards in public after finding our test set contaminated.

Your images do not become our dataset

DICOM is opened in the browser and only the pixels reach us — the header, with the patient name and ID in it, never leaves your device. The image is deleted about ten minutes after the result, a written report lives seven days, nothing is used for training, and there are no cookies and no analytics.

What it is not

Not a medical device. No CE mark, no FDA clearance, no clinical validation. It must not sit in a diagnostic pathway, and any evaluation should be read as research rather than as a second opinion.

Where this is going

An API and a workflow integration are what we are building toward, so a result can reach the place a radiologist already works instead of a browser tab. Neither exists yet, and we would rather say so than publish a page describing one. If that is the part you care about, tell us — it decides what gets built first.

Before you upload

What the result means, what happens to your scan, and how the report is opened

Six questions people ask first, answered in full — including the one that matters most, which is how to read the number you get back.