Skip to content

AI Medical Scribe Errors: Real Risks & Studies

Dr. Shahinaz Soliman, M.D. Aug 24, 2026, 4:15:15 PM
AI medical scribe error rates and studies

Contents

Quick Answer: Peer-reviewed research on AI medical scribes and clinical-note generators has repeatedly found real, sometimes serious, error rates — hallucinated exam findings, omitted symptoms, incorrect medication information — even as ambient AI adoption in healthcare grows quickly. The lesson isn't that healthcare should avoid AI. It's that AI writing into a patient's permanent record needs an architecture built around the reality that AI can be wrong: structured capture instead of free-text interpretation, and a human who remains responsible for what gets signed.

In August 2026, an Australian patient named Rebecca Green discovered something alarming in a letter from her urologist: her medical record stated she had microdosed psychedelic mushrooms. She hadn't. She'd consented to having her appointment transcribed by an AI scribe, and somewhere in that process, the system inserted a claim about illegal drug use that was never said in the room. Her doctor apologized and corrected the record, but couldn't fully explain how the error happened — only that it appeared to originate during dictation or transcription. Green, who was receiving workers' compensation at the time, said any mention of illegal drugs on her chart could have real consequences for her case (ABC News Australia, Aug 14, 2026).

It's a striking story. It's also not an outlier.

What makes documentation errors like this one different from most software bugs is permanence. A chart entry doesn't just affect the appointment where it was created — it follows the patient to every future provider who reads that record, every referral, every prior-authorization request, and in Green's case, potentially a workers' compensation claim. An AI scribe that hallucinates one line into a note isn't a one-time glitch; it's a fact that now has to be actively found and corrected before it does harm, and there's no guarantee anyone catches it before it does.

The Evidence Is No Longer Anecdotal

Over the past two years, AI scribes and ambient documentation tools have moved from pilot programs to routine use in a large share of U.S. practices. Alongside that adoption, a growing body of peer-reviewed research has quantified how often these systems get it wrong — and how serious the consequences can be when they do.

A pragmatic prospective study of 356 physician-evaluated AI-generated clinical notes found that 18% contained accidental omissions, 11.5% contained outright hallucinations — content that was never said — and 9.3% contained accidental inclusions. Most concerning: 5.3% of the notes evaluated contained errors the researchers classified as carrying serious or imminent risk to the patient (study, 2026). The researchers weren't arguing AI documentation doesn't work — overall note quality was rated high. Their conclusion was narrower and more useful: careful clinician review of AI-generated notes remains imperative, because when errors occur, some of them are dangerous.

A separate peer-reviewed evaluation of five ambient digital scribe platforms, using simulated ambulatory encounters, found a substantial mean rate of key clinical elements that were either erroneous or omitted entirely across the platforms tested — including medication information, fabricated test results, and an omitted aspirin allergy. In one simulated pneumonia-with-sepsis case, a generated note downgraded a documented plan of immediate hospitalization and antibiotics to a far less urgent one; researchers classified that specific error as carrying potential risk of death (Mayo Clinic Proceedings: Digital Health, 2026).

The American Medical Association has documented similar failure patterns directly: an ambient scribe that recorded a prostate exam had been performed when the physician had only discussed scheduling one, and another that converted a conversation about a patient's hands, feet, and mouth into a documented diagnosis of hand, foot, and mouth disease (AMA). The AMA's own guidance to physicians is direct: review AI-generated notes, diagnoses, and codes before accepting them — the physician remains responsible for what they sign, not the software (AMA).

This isn't limited to note-taking. A 2026 physician red-team study published in npj Digital Medicine had 16 physicians evaluate 888 responses from four leading AI models to real patient medical questions. Between 21.6% and 43.2% of responses were rated problematic depending on the model, and 5–13% were rated outright unsafe — with some examples researchers judged could plausibly result in serious patient harm (npj Digital Medicine, 2026). And in the UK, GP patients in South Yorkshire reported hanging up on an AI receptionist that struggled with regional accents — a reminder that even administrative, non-clinical AI needs a real path to a human when it can't do its job, not just a promise that one theoretically exists (The Guardian, Aug 20, 2026).

Consumer-facing clinical AI shows the same pattern at a larger scale. A 2026 Nature Medicine study stress-tested a consumer AI health-triage tool across hundreds of simulated clinical scenarios, including gold-standard emergency cases such as diabetic ketoacidosis and impending respiratory failure. The researchers found the system undertriaged a substantial share of those emergencies and concluded that missed high-risk cases and inconsistent safeguards raised safety concerns requiring further validation before broader consumer deployment (Nature Medicine, 2026). The takeaway isn't that consumer AI shouldn't answer health questions at all — it's that letting an AI make the call on what counts as an emergency, with no verification step, is a different and much riskier proposition than using AI to handle routine administrative work.

Why This Isn't an Argument Against AI in Healthcare

It would be easy to read all of this as a case for keeping AI out of clinical settings entirely. That's not the right conclusion, and it's not the one the researchers themselves draw. The AMA's own scribe reporting notes that most physicians using ambient AI say it saves them roughly an hour a day — hallucinations are the exception, not the norm, in that data. The pattern across these studies isn't "AI documentation fails." It's "AI documentation works well most of the time, and still needs a human backstop for the times it doesn't" — because in healthcare, the times it doesn't are the times that matter most.

At CallMyDoc, we've built our entire architecture around that same premise, for a related but different problem: not clinical note generation, but patient call handling. We're not a scribe product, and we wouldn't claim our platform would have prevented the specific documentation errors described above — those are a different failure mode in a different type of system. But the underlying design question is the same one every healthcare AI vendor has to answer honestly: what happens when the AI encounters something it shouldn't handle alone, and who's accountable when it does?

Structured Capture Instead of Free-Text Interpretation

Ambient AI scribes are built to solve an inherently open-ended problem: listen to an unstructured conversation and generate a free-text clinical summary of what happened. That's exactly the kind of task where large language models are most prone to hallucination — filling gaps with the statistically likely rather than the actually-said, because the input itself is unstructured and the model has to interpret it.

CallMyDoc's call platform is deliberately built differently, because it's solving a different problem. Rather than generating an AI interpretation of an open-ended conversation, it guides callers through structured prompts and captures what the patient states or selects — a reason for calling, a request type, an answer to a specific question — and documents that directly, categorized and timestamped, for the practice's staff and provider to act on. Where a call can't be handled by that structured path, it's routed to a person, by the patient's own choice, not by an AI's assessment of the conversation's content.

That's the practical difference between a governed system and an ungoverned one. The failure modes documented above — hallucination, scope creep, an AI filling gaps with its best guess — show up when a model is asked to freely interpret an open-ended conversation and generate its own summary. A system constrained to physician-defined rules, structured prompts, and patient-selected options isn't making that same interpretive leap, because it was never asked to. That's a different, narrower job, with a different risk profile. It's also why, across more than 29 million patient call sessions in 44 states plus DC and the U.S. Virgin Islands, CallMyDoc's platform has never lost a call, with roughly 47% of business-hour calls resolved without staff involvement — freeing staff time for the calls that genuinely need a human's judgment, not replacing that judgment with an AI's guess.

That distinction — structured capture and routing versus open-ended AI interpretation — is also why CallMyDoc has never framed itself as an "AI receptionist" that thinks for the practice. We describe what we build as clinical communication infrastructure: software that makes sure every call is captured, documented, and reaches the right person, with the AI doing the high-volume routine work and a person remaining in the loop for anything that requires actual clinical judgment. Zero breaches across eight years isn't an accident of that design — it's the point of it.

What This Means for Your Practice

Whether you're evaluating an ambient scribe, a call-automation platform, or any other clinical AI tool, the research above points to a short, practical checklist:

  • Ask what happens when the AI is wrong, not just how often it's right. Every vendor will quote you an accuracy number. Ask specifically what their error-review process catches, and who is accountable for what slips through.
  • Distinguish structured capture from free-text generation. A system that documents what a patient explicitly states or selects carries a fundamentally different risk profile than one that generates an AI summary of an open conversation.
  • Confirm there's a real, immediate path to a human — not a support ticket, a callback queue, or a chatbot loop, but an actual person the patient or physician can reach when the AI shouldn't be handling something alone.
  • Keep the physician in the review loop for anything the AI writes into the chart. The AMA's guidance applies well beyond scribes: AI-assisted documentation should be reviewed, not rubber-stamped.

None of this means healthcare should pull back from AI. The efficiency gains are real, and so is the burden AI can lift from overworked staff and physicians. It means the vendors building this software — CallMyDoc included — owe practices an honest answer about where the AI's judgment ends and a person's begins. See how CallMyDoc's hybrid AI platform is built around that boundary, or book a demo to see the architecture in action for your specialty.

Related Reading

Discover how CallMyDoc's hybrid AI approach minimizes errors and enhances patient communication. Book a demo today to see how we can help your practice.