Summary
No vendor publishes a word error rate you can compare, and if one did it would not tell you what you want to know. Accuracy is not a percentage of words. It is whether the note evidences the decision you made.
A note can be transcribed perfectly and still be wrong in the only way that costs you. From 2026 your E/M level turns on Medical Decision Making,3 so a fluent note with thin MDM is a 99214 that pays as a 99213.2
Drafted against the Medical Decision Making criteria, not just transcribed fluently.
Book a demoWhy nobody publishes a number
Word error rate is measurable and almost meaningless here. Mishearing "the" is not the same as mishearing "no". A product could score 98 percent and drop a negation, and a product could score 94 percent and never lose a clinical fact.
There is also no shared benchmark. No agreed corpus, no agreed scoring, no independent body running it. Each vendor could publish their own figure on their own data and the numbers would not be comparable, so most publish nothing. In a survey of 121 AI scribe products, 42 percent required a demo before disclosing anything at all.44Kaczmarek K, et al. Trust Me, I Might be a Medical Device, The Problem with AI Scribes. Research Square, preprint, March 2026. Survey of 121 AI scribe products.
The one thing that has been measured under randomisation is time, not accuracy. UCLA put 238 physicians across 14 specialties through roughly 72,000 encounters. Nabla cut note-writing time 9.5 percent, DAX Copilot showed no significant change, and neither reduced after-hours record time.11Lukac PJ, Turner W, Vangala S, et al. Ambient AI Scribes in Clinical Practice, A Randomized Trial. NEJM AI 2025;2(12). 238 physicians, 14 specialties, approximately 72,000 encounters. Nobody randomised accuracy, because nobody agreed what it was.
The failure that costs money
The errors that matter are not mistranscriptions. They are omissions and inventions. A comorbidity that shaped your management and did not make the note. A negation dropped, so "no chest pain" becomes "chest pain". A plausible detail the model produced that nobody said.
The first of those is also a revenue event. From 2026 your E/M level turns on number and complexity of problems, amount and complexity of data reviewed, and risk of complications.33From 2026, Medical Decision Making is the primary basis for E/M level selection. Level turns on number and complexity of problems, amount and complexity of data reviewed, and risk of complications. A note that reads beautifully and does not carry those three is a claim that gets paid one level down, and the adjudication system is right to do it.
Up to 19 percent of E/M visits are undercoded at roughly $37 each, and Medicare Advantage downcoding runs $42 to $74 a visit without ever appearing as a denial.22AAPC Audit Services reports up to 19 percent of E/M visits are undercoded nationwide at roughly $37 each. Reported Medicare Advantage downcoding runs $42 to $74 per visit and $28,000 to $74,000 per family physician per year. Word error rate does not reach any of that. See our guide to downcoding and our guide to undercoding.
The trap under the fluency
A well-written wrong note is harder to correct than a blank page. It reads as finished, so you skim it. Clinicians report that editing a long fluent inaccurate note is more draining than writing from scratch, and the more fluent the output the less carefully it gets read.
You retain full responsibility for the record regardless. NHS England states it explicitly, and outputs must be checked and corrected before saving.55NHS England guidance requires that outputs be checked and corrected before being saved to the record, with the clinician retaining full responsibility for accuracy. Additional validation is recommended for translation. No vendor in this category takes that liability off you, including us.
What to ask instead
Not the word error rate. Ask whether the note is drafted against the 2026 MDM criteria or transcribed fluently and no more, because those produce different documents and only one defends a 99214. Ask what happens to a negation. Ask whether the product will assert something nobody said, and what it does when it is unsure. Ask to run it on your own list for a fortnight, which is worth more than any number a vendor gives you.
What WA\ does
WA\ Clinician drafts against the Medical Decision Making criteria and attaches suggested coding at the point the decision was made. The comorbidity that shaped your management is in the note because it shaped your management, and the data reviewed is listed because you reviewed it.
You review and sign. The coding is a suggestion and the record is yours.
What we are not claiming
We do not publish a word error rate either, and we have just spent an article explaining why the number would be meaningless. You should read that as self-serving and test us on your own list rather than take it.
We have no randomised evidence. Nabla does, on time rather than accuracy, and it is still more than we can show you.
Availability
WA\ Clinician runs a 14-day free trial with no minimum term. Run it on real encounters and read every note it makes for a week. Pricing is on one page.
Frequently asked questions
What is the word error rate of AI medical scribes?
No vendor publishes one you can compare, and the figure would be close to meaningless anyway. Mishearing "the" is not the same as mishearing "no", so a product could score 98 percent and drop a negation while another scores 94 percent and never loses a clinical fact. There is also no shared benchmark, no agreed corpus and no independent body running one, so any published figure would be a vendor scoring itself on its own data. In a survey of 121 AI scribe products, 42 percent required a demo before disclosing anything at all.
What kind of errors do AI scribes actually make?
Omissions and inventions rather than mistranscriptions. A comorbidity that shaped your management and did not make the note. A negation dropped, so no chest pain becomes chest pain. A plausible clinical detail the model produced that nobody said. The first is also a revenue event, because from 2026 E/M level selection turns on number and complexity of problems, amount and complexity of data reviewed, and risk of complications. A note that reads beautifully without those three gets paid one level down, correctly.
Are AI scribe notes safe to sign without reading?
No, and no vendor in this category takes that liability off you, including us. NHS England states explicitly that outputs must be checked and corrected before being saved and that the clinician retains full responsibility for the accuracy of the record. There is a specific trap in fluency. A well-written wrong note reads as finished, so it gets skimmed, and clinicians report that editing a long fluent inaccurate note is more draining than writing from scratch. The more fluent the output, the less carefully it tends to be read.
Read every note it makes for a week.
Start the trial and run it on real encounters. If the notes do not hold up, you have lost a fortnight and nothing else.
About this article. Written and published by WA\, which sells an ambient documentation product and does not publish a word error rate, which you should weigh against the argument here that word error rate is the wrong measure. The randomised trial cited is independent of us and measured time rather than accuracy. Coding and downcoding figures are third-party analysis. Suggested coding is a suggestion and the clinician is responsible for the record. WA\ has no peer-reviewed clinical trials published to date and does not claim any.
