AI radiology reporting that drafts, never signs.
Fewer radiologists, more studies, every year. I build the machine that absorbs the drafting hour. Every report stays a draft until your physician signs it, and that constraint lives in code.
30 minutes. The price on this page is the price on the call.
“IMPRESSION: Stable 6 mm right upper lobe nodule. Follow-up CT in 12 months per template.”
Golden set: your historical signed reports · thresholds are contract acceptance criteria · illustrative values
The reads need a radiologist. Much of the report around the read does not.
The drafting hour is the bottleneck. The signature is not.
An imaging group’s reporting stack is rented by the seat: the dominant dictation platform runs $5,000 to $10,000 per radiologist per year, which for a 10-radiologist group is $50,000 to $100,000 every year before a single AI feature. And with the incumbent being sunset, most of the market has to re-decide its reporting stack right now. The choices on the table:
Squeeze the reading room
The attrition numbers are what that produces. Burned-out readers leave, and the residency pipeline takes years to replace them.
you lose radiologistsSubscribe to an AI reporting platform
Capable products with vendor-claimed results, no public price anywhere, and per-seat contracts. The group owns nothing when the contract ends.
no price, no ownershipLocums and teleradiology overflow
Real capacity at a premium rate, in a market that is short roughly 1,500 radiologists and heading toward 3,100.
the math gets worseA findings-to-draft pipeline, fixed scope, delivered in eight weeks
Intake
Structured findings, dictated findings, or worklist data flow in from your RIS or PACS export path. No re-keying, no workflow change for the reading room.
Template conformance
The agent maps findings onto your report templates: your sections, your required fields, your phrasing standards, your normal statements.
Draft
A complete preliminary report, findings and impression, watermarked DRAFT, with every statement traceable to the finding it came from.
The gate
A physician reviews, edits, and signs every report. Enforced in code, not in a policy PDF. The system cannot finalize a report. Draft-only is permanent and non-configurable.
Scope is deliberately narrow: one modality or study family, your top ten report templates, one reporting-system export path. That is what makes the price fixed and the timeline honest. EEG labs and diagnostic labs are the same build with a different template family; the pipeline doesn’t care whether the findings came from a CT read or an EEG review.
The eval suite is the product
The research on generative reporting is honest, so I will be too: published studies measured hallucination rates of 3 to 6% in LLM-generated radiograph, CT, and MRI reports, and most false findings in model-generated reports were clinically significant. That is not a reason to avoid drafting agents. It is the reason a drafting agent without a regression suite and a hard physician gate should never touch a diagnostic report.
A golden dataset
Built from your historical signed reports, not synthetic studies.
Template-conformance tests
Section structure, required fields, phrasing standards. A draft that breaks your template fails the suite.
Hallucination and omission regression
Every statement in a draft is checked back against the source findings, and every finding is checked for a corresponding statement. Fabrications and silent omissions both fail.
Physician-edit-rate dashboard
Instrumented from day one. If your radiologists start rewriting the drafts, the numbers say so before anyone has to complain.
A parallel-run gate
The agent drafts alongside your current process until it clears thresholds on live studies. No cutover before the score clears.
When I hand over, you own the eval harness, the golden dataset, the regression suite, and a runbook. You don’t have to trust an accuracy claim, mine or anyone’s. You re-run the suite: this quarter, next quarter, after every model update. The machine does the work, and the human judges it. In diagnostics that is not a slogan; it is the design constraint everything else hangs from.
The price is $40,000. Here is exactly what it buys.
One price. No tiers, no “starting at,” no demo-call theater. For context: a 10-radiologist group pays $50,000 to $100,000 every year for dictation seats alone. This is below the low end of one year of that, paid once, owned.
- Discovery workshop: report taxonomy, template inventory, where the sign-off gate sits
- Findings-intake pipeline from your RIS/PACS export path
- Template-conformance drafting agent (findings and impression, your formats)
- Physician review-and-sign workflow with a draft watermark until signature
- Golden dataset built from your historical signed reports
- Eval harness: template conformance, hallucination and omission regression
- Physician-edit-rate dashboard, instrumented from day one
- Guardrails in code: no auto-finalization, PHI constraints, confidence-floor escalation to a full manual read
- Parallel-run period against your current process, with published thresholds
- Handover pack: runbook, eval documentation, your team trained to run the suite
Against the alternatives
| This build | Enterprise dictation platform | AI reporting subscription | Freelancer | |
|---|---|---|---|---|
| Price | $40,000, once, public | $5,000 to $10,000 per radiologist, every year | Undisclosed; demo call required | Hourly, open-ended |
| Timeline | 8 weeks + 30-day stabilization | Months of procurement | Weeks, but never yours | Unknown |
| Who owns it | You: code, evals, golden set | The vendor | The vendor | You, minus the proof |
| Proof it works | Eval suite you re-run on your reports | n/a (dictation, not drafting) | Vendor-claimed results | None shipped |
| Physician control | Draft-only gate enforced in code | Physician dictates everything | Review-and-sign, vendor-defined | Whatever got built |
A pilot study measured a 24% reduction in median reporting time from AI draft reports without compromising accuracy. A group that reclaims even half of that recovers the build cost inside the first year. I won’t promise your number; that is what the parallel run and the edit-rate dashboard measure. I will show you the score before you cut over.
How the eight weeks run
- Weeks 1–2
Discovery + golden set
Map the reporting workflow, inventory the top ten templates, pull historical signed reports, fix acceptance thresholds in writing.
- Weeks 3–4
Intake + template engine
Stand up the RIS/PACS intake path and the template-conformance layer; first drafts against the golden set.
- Weeks 5–6
Drafting agent + the gate
Full findings-to-impression drafting, the physician review-and-sign workflow, guardrails in code.
- Weeks 7–8
Parallel run
The agent drafts alongside your current process on live studies. Cutover only when template-conformance and hallucination/omission thresholds clear.
- Days 1–30 after
Stabilization
I watch the evals and the edit rate, fix drift, tune escalation. Then you own it outright.
Straight answers
The questions every managing partner, lab director, and compliance counsel asks before booking the call.
Does the AI ever finalize or sign a report?
No, and it cannot be configured to. Every report is watermarked DRAFT until a physician signs it, and the constraint lives in code. Clinical liability is the reason diagnostic AI projects die in legal review; draft-only by design is how this one survives it.
What exactly does the $40,000 include?
Everything in the scope list above: discovery, intake, the template-conformance drafting agent, the physician gate, golden dataset, eval harness, edit-rate dashboard, parallel run, handover, and 30 days of stabilization. The only costs outside the number are your own API/hosting spend and the optional monitoring retainer.
Which modalities and report types are in scope?
One modality or study family and your top ten report templates, through one reporting-system export path. That is what keeps the price fixed. Radiology is the obvious fit, but EEG and diagnostic-lab reporting are the same pipeline with different templates. On the call we check whether your mix fits; if it doesn't, I'll tell you.
What about hallucinations?
Published rates for LLM-generated reports run 3 to 6% depending on modality, which is exactly why this build ships hallucination and omission regression, statement-level traceability back to source findings, and a physician gate that cannot be bypassed. The eval suite exists because the failure mode is real.
How do you prove it works on our reports, not a demo dataset?
The golden dataset is built from your historical signed reports, and the contract's acceptance criteria are eval thresholds on that set. Then a parallel run tests the agent on live studies, with your radiologists' edit rate tracked, before anything cuts over.
Is this HIPAA compliant? Is it a medical device?
The architecture is BAA-ready: PHI guardrails in code, minimum-necessary data handling, and the option to run on infrastructure you control. It is positioned and built as draft-assist documentation software: the physician reads the study and owns the diagnosis. The regulatory posture is reviewed with your counsel in discovery, stated plainly, not hand-waved.
What are the ongoing costs after handover?
Your own LLM API and hosting, typically $300 to $900 a month at group volume. Optionally, I keep running the eval suite and watching the edit rate for $2,000 a month, cancel anytime. You own everything either way.
Drafting is machine work. Judgment is not.
The workforce math is not turning around: fewer radiologists, more studies, every year. The drafting hour is the part a machine can absorb, and the signature is the part it never should. This build takes the first and hard-codes the second.
30 minutes · the price stays $40,000 · if it’s not a fit, I’ll say so