About

Two people building the tool we went looking for

PneumonAI is built by a small student team. Every number on this site is one we measured ourselves, on scans the models had never seen — and we publish them the way they came out.

A note from one of our founders

I built this because I once spent three days staring at a file I could not open.

I had a chest X-ray done and left the clinic with a folder of files. Not a picture — a folder, full of DICOM, the format hospitals actually use. My laptop offered to open it with a text editor. The clinic’s own reading would come in a few days, and until then there was an image of the inside of my chest sitting on my desk, and I had no way to look at it.

So I did what anyone does. Free viewers that wanted an install and a licence key. Forums full of people asking my exact question and being told to see a doctor — true, and useless at eleven at night. And a scattering of AI demos that would take a JPEG and hand back a number to two decimal places, with no indication of what it measured, what it had been trained on, or how often it was wrong. The tools that were honest told me nothing. The tools that told me something were not honest.

It turned out fine. But those three days stayed with me, and so did the exact shape of the problem: the file existed, the models that could read it existed, and none of it was reachable by the person whose chest it was.

That is the whole reason this site exists, and you can see it in the decisions. DICOM opens here, in your browser, with the window level the clinic stored inside it — because that was the first wall I hit. There is a written report and not just a verdict, because a number on its own was exactly what I already had too much of. And every figure we publish comes with the images it was measured on, including the time our own first result came out too flattering and we corrected it downwards in public.

We are not trying to replace the person who reads your scan. We are trying to make the wait less blind.

Arsenii Ahamalov · co-founder

Mission

Make an AI reading of a chest X-ray something that can be checked rather than believed. For a department that means every measurement we hold, on scans the models never saw, published whether or not it flatters us. For a person holding their own result it means the same facts in language that does not need a medical degree.

Team

ML & Frontend

Arsenii Ahamalov

LinkedIn

Backend & C++ Inference Engine

Artem Romanov

LinkedIn

Technology

Models

  • PyTorch — training; TorchScript for everything that ships
  • TorchXRayVision — the three pneumonia classifiers and the lung segmentation model
  • segmentation-models-pytorch — the localizer, the one model we trained ourselves

Inference

  • C++ / libtorch — the pipeline: decode, validate, segment, classify, localize — CPU only
  • Redis · PostgreSQL — jobs live ten minutes; reports live seven days

Explanation

  • Claude Sonnet 5 — the written report and the follow-up chat, from the measurements only — it never sees the image
  • FastAPI · WeasyPrint — the report service and the PDF you can hand to a doctor

This site

  • React · TanStack Start — server-rendered, no analytics, no cookies
  • Docker · nginx — one machine, one compose file

How this got here

PneumonAI started as a single model and became an ensemble of three DenseNet121 classifiers, trained independently at Stanford, at MIT and Harvard, and at a university hospital in Alicante. Three institutions rather than one, because models that fail in the same places agree with each other and tell you nothing.

August 2026 · a member removed

One of the original three came from the NIH ChestX-ray14 dataset. Its vote was sound — on paper it was the most accurate member we had. Its heatmap was not: across 25 000 chest X-rays the point it highlighted landed on the border of the image 91% of the time, and inside the lung fields in 4%. The cause was not a bug we could fix. That model only calls a scan positive at a very low score, and the arithmetic behind the heatmap discards everything below zero, so what was left to display was the emptiest corner of the frame.

We replaced it with a model trained on PadChest. On data neither model had seen, it also scored higher than the one it replaced — the earlier model's apparent edge came from being tested on images from its own training hospital. The decision threshold was re-measured afterwards, because a threshold belongs to an exact set of models and stops meaning anything the moment one is swapped.

August 2026 · the heatmaps went too

Having fixed one model's attention, we measured all of them properly, against regions drawn by radiologists. The answer was worse than the fix: the surviving classifiers' heatmaps put their strongest point inside a marked region only 1.3 to 2.7 times more often than a point placed at random. Three maps disagreeing under one verdict did not help a reader pick the right one — it undermined trust in all three at once.

So we removed the mechanism entirely and trained a model whose only job is to answer where. On held-out scans its peak lands inside a marked region 86% of the time — 8.1 times chance. It has no vote on whether, and on 18% of positive scans it marks nothing at all and says so.

We keep the removals in this history rather than quietly dropping them, because they are the honest shape of the work: a model can be right about whether and unable to say where, and finding that out cost us a member of the ensemble and an entire feature. Showing where is half of what this tool is for.