Summary
Every honest answer to this question starts with a number. Across simulated ambulatory encounters, ambient scribe platforms hallucinated clinical content at roughly 7 percent on average, and the physical examination was the most vulnerable section.
A hallucination here means the note asserts something that was not said or not done. A finding that was never examined, a symptom that was never mentioned, a value that was never given. In a clinical note that is not a typo, it is a risk.
The evidence also points to the fix. A better model on its own does not close the gap. The clinician reading and signing the note does. That is the part every serious body treats as required rather than optional.
WA\ shows its working and holds every patient-affecting action for human approval. Ask to see it on a call.
Book a demoThe 7 percent, in context
The figure comes from work evaluating ambient digital scribe platforms against simulated encounters, where the ground truth is known. On average about 7 percent of generated content was inaccurate, and the physical examination carried the highest rate because it is the section a scribe is most tempted to infer rather than record.
Directionally this holds across the category. It is not a reason to avoid scribes. It is a reason to design for review and to be honest about where the errors cluster.
What the one randomised trial found
Ambient scribes have a single randomised trial. It cut note-writing time by 9.5 percent, about 41 seconds an encounter, across roughly 72,000 encounters and 238 physicians in 14 specialties. Neither arm reduced after-hours time in the record.
The trial measured time and workload rather than hallucination rate directly, which is why the simulated-encounter work matters alongside it. Read together they say the same thing. The tool helps, and the note still needs a human.
The upside is also measured
The case for scribes is real and it is documented. A study of clinicians using an ambient scribe reported burnout falling from 51.9 percent to 38.8 percent after 30 days, a drop of 13.9 percentage points.
So the honest picture is two numbers held together. A meaningful reduction in burnout, and a hallucination rate that makes review non-negotiable. A vendor that quotes only the first is not giving you the whole thing.
Why sign-off is the safeguard
The peer-reviewed literature is consistent. Clinician review, correction and final sign-off are treated as essential safeguards, not best-practice extras. The scribe drafts, the clinician decides.
This is a design question as much as a policy one. A note you can check quickly gets checked. A note that hides its sources or buries its reasoning gets rubber-stamped, which is where the 7 percent becomes a real error in a real record.
How WA\ is built for this
WA\ shows its working. Notes, letters and care plans are drafted with their cited evidence attached, so the clinician is reviewing reasoning rather than a black box. Every AI action that affects a patient is reviewable and held for human approval.
We do not claim a zero hallucination rate, because no honest vendor can. We claim a design that makes the errors easy to catch before anything is signed. That is the part that actually protects the patient.
Frequently asked questions
How often do AI medical scribes hallucinate?
Across simulated ambulatory encounters, ambient scribe platforms produced inaccurate content at roughly 7 percent on average, with the physical examination the most affected section. The rate varies by product and by how it is measured, and it is why clinician review before sign-off is treated as essential rather than optional across the peer-reviewed literature.
Are AI scribes safe to use in practice?
They are safe when they are used as designed, which means the clinician reads and signs every note. The evidence shows a real benefit, including a reported 13.9 point drop in burnout after 30 days, alongside a hallucination rate near 7 percent that makes review non-negotiable. The safe pattern is the scribe drafts and the clinician decides. WA\ is built so every patient-affecting action is held for human approval.
Which part of the note is most likely to be wrong?
The physical examination. It is the section a scribe is most tempted to infer from context rather than record from what was actually said and done, so it carries the highest hallucination rate in the simulated-encounter evidence. It is the first place a reviewing clinician should look before signing.
Evidence you can check, not claims.
Thirty minutes and you will see how WA\ shows its working and holds every patient-affecting action for human approval.
