What AI Medical Scribes Get Wrong and Human Scribes Get Right

A hallucinated line in a customer service chat is an inconvenience. A hallucinated line in a chest pain note is a malpractice liability. That is the difference between physician practice owners and compliance administrators need to evaluate AI medical scribes on clinical evidence.

AI medical scribes are being adopted faster than they are being validated. 

The research is now catching up. It shows how AI-generated clinical notes contain fabricated content at measurable rates and score lower than human-written notes across documented quality. 

They perform less accurately for patients who speak African American English or non-standard dialects. 

They create HIPAA and consent exposure most practices have not resolved. Documentation gaps flow directly into billing denials, which are unignorable.  

Trained human medical scribes prevent each of these failure modes. This blog explains where AI medical scribes fail, the clinical and financial consequences you might face, and what a trained human scribe does instead (based on peer-reviewed evidence). 

AI Medical Scribes Hallucinate Clinical Content at a Rate No Practice Can Ignore 

Hallucination is one of the most serious AI scribe failures. It is not just minor typos or grammatical glitches. They are fabricated clinical data points such as fake heart sounds, incorrect lab values, or unperformed physical exams. 

An AI Medical Scribe hallucination occurs when an automated medical tool invents false medical facts, diagnoses, or treatments that were never said during a patient visit. 

How hallucination differs from transcription error and why the distinction matters clinically 

A transcription error is a mishearing. Transcription errors occur when an Automatic Speech Recognition (ASR) system misinterprets acoustic patterns. It happens due to poor audio quality, heavy accents, background noise, or homophones.  

A hallucination is different. The AI generates clinical content that was never spoken at all. This distinction matters because hallucinated content cannot be distinguished from accurate content. 

So when an ASR system mishears “denies chest pain” as “describes chest pain”, the error is traceable to audio input. When an AI model generates a documented physical example finding for a test that was never performed, there is no audio to check. 

Result: Fabrication exists only in the note that is written in the clinician’s expected tone, format, and positioned where content would appear. 

The specific content categories where hallucination appears most often 

Hallucinated content appears most often in the sections of the clinical note where AI language models are most likely to generate “plausible completions”. 

Common examples include physical exam findings, assessment and plan, medication lists, and documented notes like “denied fevers” or “no history of cardiac disease”. 

These are high stakes as they inform about diagnosis, guide treatment, and support medical necessity for billing. 

A validated blinded study using the PDQI-9 framework found hallucinations in 31% of ambient AI-generated notes compared to 20% of physician-written notes. AI introduces fabricated clinical content at a higher rate than clinicians themselves do. 

The hallucinated entries look medically reasonable, which makes it harder to catch.

Why physician review alone does not reliably catch AI-generated fabrication 

A physician reviewing AI-generated notes cannot find fabrication right away under clinical time pressure. They are working against a specific cognitive disadvantage: notes are correctly formatted. They have appropriate medical terminology and follow the expected documentation structure. 

There is no visual signal for them to highlight that a line is fabricated. This is what makes hallucinations hard to catch. 

Research on cognitive load in clinical hours consistently shows reduced error detection. A physician signing twenty notes at end-of-shift is not reading each line with scrutiny during a chart audit. 

AI-generated hallucinations are not designed to look wrong. They are designed, by the model’s architecture, to look exactly right. 

What happens downstream when a hallucinated entry passes into the permanent record 

Once a hallucinated entry is signed, it becomes part of the permanent medical record. It may appear in future visit notes. A subsequent provider reading that note may document a response to a finding that never existed. 

A coder may select an ICD-10 code supported by a hallucinated diagnosis. Similarly, a payer may approve or deny a claim based on documentation that does not reflect what occurred in the encounter. Healthcare documentation errors continue to affect patient care and billing long after the original visit.

It can cause legal risk during a medical review, where the medical record is treated as evidence in malpractice proceedings.

A hallucinated physical exam finding that was never caught, never corrected, and never flagged becomes part of the factual record as the court reviews. But the AI vendors bear no liability for that entry. The physician who signed it does.

AI Medical Scribes Produce Lower-Quality Notes Across Every Measured Documentation Domain 

AI medical scribes can produce a medical note within seconds. But speed doesn’t ensure accuracy. 

What the PDQI-9 measures and why it is the clinical standard for note quality evaluation 

The Physician Documentation Quality Instrument (PDQI-9) is a validated, blinded evaluation framework used in peer-reviewed research to measure the quality of clinical notes across ten defined domains. 

These domains include: accuracy, thoroughness, usefulness, organization, comprehensibility, conciseness, synthesis, internal consistency, and whether the note reflects as the current and relevant clinical information.

Each domain is rated on a five-point scale from 1 (not at all) to 5 (extremely) by reviewers who do not know whether the note was written by a clinician or generated by an AI tool.

This blinded design is what makes the PDQI-9 framework findings credible for the comparison. Reviewers cannot favor human over AI notes. 

They evaluate the notes on clinical merit alone. When multiple independent studies, using this instrument, find the same result, it carries evidential weight. 

How AI note scores compare to human clinician scores across accuracy, thoroughness, and usefulness 

A cross-sectional evaluation published in the Annals of Internal Medicine compared AI-generated and human clinician notes using the modified PDQI-9 framework. 

AI scribe notes scored lower than human clinician notes across all ten documented quality domains. The largest performance gaps appeared in thoroughness, usefulness, and organization. 

The usefulness finding carries one significant operational implication: a note can be technically complete, contain accurate information, but still fail to support clinical decision-making. 

A note that does not communicate the clinician’s reasoning is not a usable note. It does not matter how quickly it was produced. 

A separate Mayo Clinic Emergency Department pilot evaluating 710 patient encounters (December 2024 to January 2025, published in Annals of Emergency Medicine) found that human scribes produced higher-quality notes in pediatric encounters as compared to AI scribes. 

The quality gap, in other words, has a direct revenue impact. 

These issues can even become harder to manage when information moves between different systems, creating documentation gaps from EHR interoperability failures. When key details are missing or disconnected, clinicians may have to make decisions without the full picture of the patient’s history. 

Where specialty context changes the quality gap most sharply 

The quality gap between AI and human scribes is not uniform across visit types. Research consistently identifies where AI documentation performance declines: multi-problem visits, pediatric encounters, psychiatric consultations, oncology follow-ups, and high acuity emergency medicine. 

For a simple visit, such as treating a sore throat or a minor infection, AI scribe tools can produce an adequate structural note. Note complexity is low enough that it poses limited risk. 

But a complex visit, such as managing diabetes, high blood pressure, and several medications at the same time, gives it much more to document correctly.

Speed and note quality are not the same things. An AI scribe may create a note in eight seconds. If the doctor has to check the medical history, fix the assessment, correct the treatment plan, and rewrite parts, the time that was saved using an AI medical scribe is lost. 

AI Scribes Produce Systematically Less Accurate Notes for Specific Patient Populations

AI scribes don’t understand every patient’s speech equally. Heavy accents or poor audio quality can cause mistakes in taking notes. These mistakes create issues for patient care. 

Why automatic speech recognition error rates are not uniform across patients 

AI scribes use speech-to-text technology as their first processing step to convert spoken audio to text. But this technology does not understand every patient’s speech equally, making it a variable in AI scribe documentation quality. 

ASR systems are trained on audio datasets. The composition of those datasets determines which speech patterns the system handles most accurately. For instance, commercial ASR systems are trained on standardized American English speech. So they perform less accurately on regional accents and speech patterns. 

It is a documented disparity confirmed across multiple peer-reviewed evaluations of commercial ASR systems. 

How ASR bias against African American speakers and non-standard accents translates into clinical documentation risk 

A study published in JAMIA Open (Zolnoori et al., 2024) found that commercial and clinical ASR systems consistently showed lower transcription accuracy for Black patients compared to White patients. 

A separate evaluation of five major commercial ASR systems (Koenecke et al.) found that error rates for African American English speakers were nearly double those for White speakers.  It was recorded across all five systems evaluated, not from a single platform. 

For example, a patient may say, “my pain has gotten worse since last week, especially at night.” If the system misses the word “a lot worse” or “especially at night,” the note may document a stable or mildly progressing condition. 

This can be especially risky in medical visits where patient symptoms guide the diagnosis and treatment. If important details are missed, the doctor may have an incomplete picture of the patient’s condition.

What the documentation gap means for patients from underserved communities 

The problem is that ASR accuracy disparity does not affect all patients equally. 

Patients who speak African American English, patients whose first language is not English, and patients from rural or regional communities with distinct speech patterns are systematically more likely to receive AI-generated clinical notes. It will easily misrepresent what they said. 

But a trained human scribe brings capabilities to this gap that no current ASR model can replicate. 

As documented in a PMC commentary (Topaz, Zhang, and Peltonen, npj Digital Medicine, 2025), human scribes from shared community backgrounds can recognize dialect-specific phrasing, interpret emotional context of how something is said, and document social determinants of health that are embedded in how patients speak. 

But an AI system processing an audio input cannot observe hesitation or recognize cultural phrases that a patient may ask to clarify a question.

This difference matters as the medical record is used in many places for patient treatment, and the next provider doesn’t know the medical notes have errors in them. 

AI Scribes Create HIPAA and Consent Exposure That Human Scribes Do Not 

Using an AI scribe is more than writing medical notes. It can also create new privacy and consent risks because patient information may be recorded, processed, or shared through the AI system. 

Why ambient recording creates a different HIPAA risk profile than human scribe documentation 

A human scribe follows HIPAA foundations like defined BAA, physical access protocols, and documentation that remains within the practice’s own EHR. There is no ambient recording, no third-party cloud storage or API connection to an external vendor. 

An AI tool operates differently. It records the conversation. Audio is transmitted to the vendor’s cloud infrastructure for processing. The processed text is returned via an API connection to the practice`s EHR.

Depending on the data retention and model training policies, the audio or transcript may be retained and used to improve the AI system. Use of PHI is prohibited by practices. Most patients have not explicitly consented to use of their medical information. 

According to legal analysis published by Foley and Lardner LLP (September 2025), it can create a compliance obligation. A practice that has not mapped every PHI touchpoint in its AI scribe implementation has not completed its HIPAA Security Risk Assessment. 

The consent complexity AI scribes introduce that most practices have not resolved

HIPAA and patient consent are separate legal frameworks. AI scribe adoption created obligations under both. 

A signed BAA covers HIPAA requirements. But it does not satisfy the patient’s right to know their visit is being recorded or transmitted to a third-party vendor. 

Eleven U.S. states have two-party or all-party consent laws that require all participants to have explicitly consented to the recording. 

How AI scribe data storage and vendor training practices expand the PHI exposure surface 

Before deploying an AI medical scribe, practices should review their HIPAA security Risk Assessment that maps every new PHI data flow, have a BAA with the AI vendor when required, and have a written policy prohibiting clinicians from using personal devices for AI scribe recording.  

The bigger issue is that “HIPAA compliant” does not mean the practice can stop asking questions. The practice still needs to understand what the vendor will do with patient data and where it will be stored. Who can access it, how long they will keep it, and whether it’s used for other purposes. 

Documentation Errors From AI Scribes Flow Directly Into Billing Denials and Revenue Loss 

Thin or disorganized notes don’t just create clinical risk; they directly shape whether a claim gets paid, queried, or denied, making documentation quality a revenue cycle issue as much as a clinical one. 

How incomplete or inaccurate AI-generated notes translate into claim denials 

Documentation quality and claim integrity are directly connected. Every note helps coders choose the right ICD-10 and CPT codes and shows why a particular medical service was necessary. If important details are missing, it results in a coder query, a downcode, or a payer denial. 

It means documentation quality affects more than patient care. It can also affect how accurately a practice gets paid. The connection between documentation and payment is why practices need to understand how documentation quality connects to revenue cycle outcomes. Incomplete or inaccurate notes can lead to coding errors, claim denials, and delays in reimbursement. 

The specific documentation elements AI scribes miss that coders and payers require 

AI scribes can transcribe what physicians say but still miss clinical reasoning that payers require to process a claim. 

A physician may verbally document that a procedure was performed, but if the document does not explain why the procedure was necessary, the clinical indicators that support the decision will be missing. A claim may be denied for lack of medical necessity documentation. 

Why specialty practices face higher revenue risk than primary care 

Specialty practices, such as cardiology, orthopedics, oncology, neurology, and psychiatry, operate under coding frameworks. These practices require procedure-specific documentation language. 

Both coders and insurance companies need specific details in each specialty-related note. Words like ‘severe’ or ‘worse’ may not be enough. AI scribes are not trained on specialty-specific coding requirements. 

They transcribe clinical language. 

A trained human scribe with specialty experience understands which clinical details cardiology coder needs to support a specific CPT code. 

This is why the specialty has higher documentation risk and a wider gap between what an AI scribe captures. 

Trained Human Medical Scribes Outperform AI in the Clinical Encounters Where the Stakes Are Highest 

AI scribes can struggle with complex, high-risk medical visits. Trained human scribes can catch missing details, ask for clarification, and make sure the note accurately reflects the visit. For complex care, practices may still need trained human scribes when AI alone is not enough. 

Finding qualified healthcare staff for these roles can be difficult, especially when practices need people with experience in specific clinical settings. Healthcare staffing shortages can make it even harder to build that support team. This can leave practices with staffing gaps that slow down workflow and put more pressure on existing clinical teams. 

The specialty and encounter types where AI scribe failure risk is highest 

AI scribes can struggle more with complex visits, especially those involving multiple health problems, pediatric care, psychiatric visits, oncology, and emergency medicine. These visits often involve more details, which makes the documentation process even more complicated. 

What human scribes capture that AI cannot: nonverbal cues, emotional context, social determinants 

A trained human medical scribe is not a transcription device. 

In a complex evaluation, even a virtual human medical scribe is an observational participant. They are present in the room, watching the patient, tracking interaction, and documenting not only what is said but also what is observable, like the patient’s tone, pace, body language, or visible distress. 

An AI scribe cannot make any of these observations. It works based on recorded conversations, so it may miss details that are visible but never spoken. 

How the hybrid model works and why the human scribe role in it is not optional in complex care 

A hybrid model doesn’t mean AI handles every note and the human scribe only checks medical notes. 

A hybrid model works like this: AI documentation tools handle routine, low-risk visits so their output can be reliable. Trained human virtual medical scribes handle complex specialty visits like psychiatric consultations or oncology follow-ups, where accurate documentation matters most, and AI cannot be trusted, or financial stakes of documentation failure are high. 

However, staffing that human scribe capacity has its own challenge. Building specialty-trained documentation support for complex care requires a recruitment strategy that covers both clinical training requirements and operational realities. 

For practices that need human support for complex documentation, trained virtual medical scribes available in less than 7 days offer a structured path to human documentation support without the delays of traditional hiring. 

Conclusion

AI medical scribes fail in specific, documented ways: they hallucinate clinical documents at measurable rates, score lower than human-written notes, perform less accurately for patients who speak non-standard dialects, create HIPAA exposure, and produce documentation gaps that translate into billing denials.

These are the findings from peer-reviewed research and institutions with no financial interest in the outcome of the comparison. 

For routine, low-complexity visits, AI documentation tools can reduce administrative burden. But where clinical accuracy, billing integrity, and patient safety are most at stake, trained virtual medical scribes remain the documented standard. Hence, they remain valued by specialty practices that cannot afford billing risks or losing their patients. 

Most Frequently Asked Questions

If an AI scribe hallucinates something in a note and the physician signs it, who is legally liable?

The physician who signs the note bears legal responsibility for its accuracy, regardless of whether AI generated the content. Malpractice defense for AI-related documentation errors is increasingly framed around “failure to supervise AI output”. It means the physician is expected to verify AI-generated content before signature.

Yes, AI scribes can process accented speech, but peer-reviewed research documents that accuracy is measurably lower for non-standard accents and dialects. A study published in JAMIA Open (Zolnoori et al., 2024) found lower transcription accuracy for Black patients across commercial and clinical ASR systems. Separate research (Koenecke et al.) found error rates for African American English speakers nearly double those for White speakers across five major commercial platforms. 

For complex encounters, yes, and evidence supports it. A Mayo Clinic Emergency Department pilot study found that physicians contributed more to AI-generated notes than to notes prepared by human scribes, suggesting more editing was needed. AI can save time for routine visits, but complex medical visits still require a proper physician review and editing. 

A practice should verify four things: 

  • The vendor will sign a Business Associate Agreement (BAA).
  • Confirm your practice has completed an updated HIPAA Security Risk Assessment.
  • Obtain written answers from the vendor on whether patient audio is stored or used to train AI models.
  • Establish and document a practice policy prohibiting clear rules against recording patient visits on personal devices.

To be HIPAA-eligible doesn’t mean HIPAA-compliant. Practices should verify the vendor’s actual privacy and security practices before using the tool.

There’s no current evidence to support this conclusion. A cross-sectional study published in the Annals of Internal Medicine (Reddy et al., April 2026, VA Puget Sound) found that AI scribe notes scored lower than human clinician notes across all 10 PDQI-9 quality domains under blinded review conditions. It had the largest gaps in thoroughness, usefulness, and organization. AI adoption is moving faster than clinical validation to confirm safety and accuracy. 

AI scribes work more reliably in telehealth than in complex in-person visits. The major difference is the environment. ASR performance is more consistent in telehealth. In-person visits introduce multiple variables that lower ASR accuracy, such as background noise, multiple simultaneous speakers, patient emotional distress, and physical examination sounds. 

A virtual human scribe is a trained clinical documentation professional working remotely in real-time. They carry the same clinical judgement, contextual interpretation capability, and compliance footprint as an in-person human scribe. They are trained to use EMR/EHR platforms and keep every detail in check. 

An AI scribe is an automated software that processes ambient audio and generates a note without human judgment. The two should not be treated as the same because they have different accuracy, oversight, and data-handling risks.

Subscribe to Our Newsletter
Receive occasional updates, hiring insights, and practical tips on building reliable remote teams, sent only when it’s useful.

Build Your Expert Remote Team in Less Than 10 Days.
Hiring top-tier talent is simple, fast, and reliable through Remote Scouts. 100% risk-free virtual assistant staffing with top 3% vetted candidates across multiple industries and regions. No more work delays.

"*" indicates required fields

This field is for validation purposes and should be left unchanged.
Begin Your Risk-Free Hiring Process
Looking for a job? View Our Current Openings.