Two years ago, an ambient AI scribe was something a health system piloted with twenty volunteers. In 2026 it is a line item in a three-physician practice's software budget. The American Medical Association's 2026 Physician Survey on Augmented Intelligence, published March 12 and based on 1,692 physicians surveyed in January and February, found that 81 percent of physicians use AI in their work, up from 66 percent in 2024 and 38 percent in 2023. Documentation of clinical care and summarizing research lead the use cases, and physicians reported an average of 2.3 uses each, up from 1.1 three years ago. The share who describe themselves as uncertain about AI fell to 9 percent, from 18 percent two years earlier.

From the billing side, this is a bigger change than it looks. The note is the evidence for the claim. When the person writing the note changes from the physician to a model listening to the room, the evidence changes too, in ways that are mostly good and occasionally expensive. We have now reviewed enough scribe-generated notes across specialties to have a short list of what goes wrong.

Key takeaways

  • Scribe notes capture history well and get signed faster. That alone removes a major source of unbilled encounters.
  • They often describe medical decision making without deciding it. The physician has to state the assessment and plan, aloud or in an edit.
  • Automatic time statements based on recording length are not the physician's total time. Turn them off.
  • Problems the patient mentioned but the physician did not address inflate the problem count. That is upcoding, intended or not.
  • Audit 20 notes per provider after 60 days, blind, and compare the level distribution to the year before.

What gets better

Let us be fair to the technology. Scribe notes almost always capture the history of present illness in more detail than a physician typing between patients. They record what the patient said about medication adherence, social factors and symptom timing. They rarely leave the assessment blank. Providers who used to sign notes three days late sign them the same afternoon, which alone removes one of the largest sources of unbilled encounters. A coder reading a scribe note usually has more to work with, not less.

The AMA survey found that 70 percent of physicians see AI as a way to automate the tasks that contribute to burnout, and documentation is the task they name first. Our experience matches that. A practice that moves from three-day signature lag to same-day signature sees its charge lag fall, its held-claim queue shrink, and its month-end close move up by a week. Those are revenue cycle wins, and they are real.

The five failures we see

1. Medical decision making is described but not decided

Since 2021, office E/M levels (99202 to 99215) are chosen by medical decision making or by total time. MDM has three elements: the number and complexity of problems, the data reviewed and analyzed, and the risk of management. A scribe captures the conversation, and the conversation often does not contain the physician's reasoning. The note will say the patient has hypertension, diabetes and knee pain, list the medications, and record that labs were ordered. It will not say that the diabetes is inadequately controlled and the plan is to increase metformin, which is what makes it a 99214 rather than a 99213. Coders cannot infer that. The physician has to add it, and the practices that do well train providers to state the assessment aloud at the end of the visit.

2. Time statements that are not the physician's

Some scribe products insert a time statement automatically, based on the recording length. Total time for E/M purposes is the physician's time on the date of the encounter, including pre-visit chart review and post-visit documentation, and excluding time spent by staff. A recording length is not that number. A note that says "Total time spent: 27 minutes" generated by software, without the physician confirming it, is a documentation risk, and it can push a visit to a level the MDM does not support. Turn the automatic statement off or require the physician to edit it.

3. Problems the patient mentioned but the physician did not address

A patient says, in passing, that their back has been bothering them and their sleep is poor. The scribe dutifully adds both to the problem list in the note. The physician addressed neither. Now the note appears to show four problems addressed when two were. In an audit, that is upcoding, even though nobody intended it. The fix is a review habit: the physician deletes what was not addressed before signing.

4. Copy-forward by another name

Scribes that pull prior notes into the draft can reproduce last visit's exam and plan word for word. Payers have been flagging cloned documentation for a decade. A scribe does not make cloned text more defensible; it makes it more common.

5. Procedures in the room that never became charges

The scribe records the visit. It does not, in most products, create a charge for the joint injection or the ear lavage that happened during it. Practices that relied on the physician clicking the procedure in the old workflow sometimes lose those charges when the workflow changes. Run the encounter-to-charge reconciliation weekly for the first three months after a scribe goes live.

A worked example of the level shift

Consider an internist whose established-visit mix in 2025 was 35 percent 99213, 55 percent 99214 and 10 percent 99215. Two months after a scribe goes live, the billed mix is 20 percent 99213, 60 percent 99214 and 20 percent 99215. At Medicare rates the difference between a 99213 and a 99214 is roughly $40 to $45, and between a 99214 and a 99215 roughly $50 to $55, so on 400 established visits a month the shift is worth about $5,000 a month more.

Level2025 mixPost-scribe billed mixPost-scribe audited mix
9921335%20%27%
9921455%60%60%
9921510%20%13%

The blind audit put about half of the shift back. The 99215s that held up were visits where the physician documented high-risk prescription management and the data reviewed. The ones that did not were visits with an automatic time statement and a problem list padded with things the patient mentioned. That is the typical pattern: half the lift is legitimate, because the documentation finally supports the work, and half is not. The audit tells you which half you have. Do it before a payer does.

Setup questions before go-live

QuestionWhy it matters for billing and compliance
Is there a signed business associate agreement with the vendor?The vendor processes protected health information; a BAA is required
Where is the audio stored, for how long, and can it be produced for an audit?Some payers and auditors will ask what the recording supports
Does the state require patient consent to record?Several states require all-party consent; a scripted consent at check-in avoids the problem
Does the note carry an attestation that the physician reviewed and edited it?Most payers expect the signing provider to attest to AI-assisted documentation
Is the automatic time statement off?See failure number two
Does the vendor appear in your risk analysis?It is a new system holding ePHI; OCR expects it inventoried

A short coding audit after 60 days

Pull 20 notes per provider from the second month after go-live and code them blind. Compare the audited level to the billed level and to the same provider's level distribution from the previous year. Report by provider and by failure type, using the five above as the categories. Then hold a fifteen-minute conversation with each provider using three of their own notes: one that was right, one where the MDM was missing, one where the problem list was padded. Providers change habits when they see their own notes; they do not change them for a slide deck.

If you want an outside review, our RCM audit includes a documentation and E/M level sample, and our medical coding team works with scribe-generated notes daily.

Questions we hear

Should the coder see the recording?

No. The note is the record. If the note does not support the level, the physician amends the note with an addendum dated and signed; the coder does not code from audio.

Our providers say the scribe made their notes "too long." Is that a billing problem?

Length is not a problem. Irrelevance is. A long note with a clear assessment and plan is fine. A long note in which the assessment is buried under three paragraphs of transcribed conversation slows the coder down and hides the MDM. Ask the vendor about output templates; most can be configured to put the assessment and plan at the top.

Do we need to tell payers we use a scribe?

There is no general requirement to notify payers, but the signing provider's attestation should make clear that the note was reviewed and edited by the provider. If a payer audit asks how the note was produced, answer directly; the attestation and the audio retention policy are what they will want to see.

What to do this month

  1. Confirm a signed BAA with the scribe vendor and add the vendor to the risk analysis inventory.
  2. Turn off automatic time statements, or configure them to require provider confirmation.
  3. Add a one-line consent script at check-in if your state requires it.
  4. Run the encounter-to-charge reconciliation weekly for the first three months after go-live.
  5. Schedule the 20-note blind audit for the second month, and the provider conversations for the week after.