A practice manager told us last month that her lead biller had started drafting appeal letters with ChatGPT and that the letters were better than the ones the practice had been sending. She was pleased until we asked what the biller had pasted in. The answer was the denial letter, with the patient's name, date of birth, member ID and diagnosis. That is a disclosure of protected health information to a vendor with no business associate agreement, and it happened in a practice that runs annual HIPAA training.

ChatGPT became publicly available on November 30, 2022. GPT-4 followed on March 14, 2023. In April, Epic and Microsoft announced plans to bring the same underlying model into Epic's tools, starting with drafting replies to patient messages, and several documentation vendors have announced ambient note-writing products built on it. In the billing office, none of that has arrived yet. What has arrived is a free website that writes fluent English, and staff who have found it.

Key takeaways

  • Generative AI tools are good at form (drafting, summarizing, rewording) and unreliable at fact (codes, payer rules, dates, citations). Everything factual has to come from you or be checked by you.
  • Pasting a denial letter with patient identifiers into a public chatbot is a disclosure to a third party without a BAA. The rule staff can follow is: nothing identifiable goes in.
  • Do not let these tools assign codes or decide coverage. The vendor claims about autonomous coding are ahead of what the tools can do in 2023.
  • A three-rule policy (no PHI, a human owns the output, no coding or eligibility decisions) plus a short list of approved uses is enough for this year.
  • The time saved is writing time. Judgment, follow-up and payer knowledge stay with the biller.

What it does well today

These tools are language tools. Where the task is turning facts you supply into clear prose, they are good, sometimes very good.

  • Appeal letter drafts. Give it the denial reason, the payer policy language, the clinical facts with identifiers removed, and the argument you want to make, and it will produce a structured letter in seconds. A biller who writes three appeals a day can review and edit far faster than write from scratch.
  • Summarizing payer policy text. Paste a twelve-page medical policy and ask which criteria apply to a specific CPT code. It is a fast first read. It is not the final read.
  • Patient-facing explanations. Rewriting a statement insert or a financial policy at a sixth-grade reading level, or translating it, is a task it handles well, and a human still checks it.
  • Training and onboarding. "Explain what CO-197 means and what a biller should check first" produces a decent answer for a new hire, and the trainer can correct anything wrong.
  • Spreadsheet help. Formulas, pivot table steps, and cleaning up a payer's export are the sort of question staff used to search the web for.

Where it fails, and fails confidently

TaskOur assessment in June 2023
Assigning CPT or ICD-10-CM codes from a noteNot reliable. It produces plausible codes, sometimes invalid ones, and cannot apply payer edits or the Official Guidelines consistently. Do not code from it.
Payer-specific rulesIt does not know your contract or the plan's current policy, and it will state a rule as fact anyway.
Anything with a dateThe consumer tools were trained on data with a cutoff, so the 2023 E/M changes, the April ICD-10-CM update and the end of the PHE are unknown or wrong.
Legal or regulatory citationsIt invents citations that look real. Verify every one against the source.
Arithmetic on remittancesUnreliable for anything beyond simple sums. Use the spreadsheet.

The pattern is that the tool is excellent at form and unreliable at fact. Everything factual in its output has to come from you or be checked by you. Honestly, we think the coding claims being made by some vendors this year run well ahead of what the tools can do, and a practice that lets a chatbot pick codes will find out in an audit.

Try it yourself with a de-identified hospital admission note and ask for the E/M code. In our experience the explanation that comes back is built on the 2022 rules (history, exam and MDM as three scored components), which have not applied to hospital codes since January 1, and the observation codes it offers as alternatives were deleted the same day. A coder who trusted that answer would bill a deleted code with an outdated rationale. The tool is not lying. It simply does not know the calendar.

The privacy problem

The consumer versions of these tools do not offer a business associate agreement. By default, what you type may be retained and used to improve the model. Under HIPAA, entering a patient's identifiable information into such a tool is a disclosure to a third party without a BAA, and it is reportable if it meets the breach definition. The fix is not to ban the tool; staff will use it on their phones. The fix is a rule they can follow: nothing identifiable goes in. No names, dates of birth, member IDs, claim numbers, dates of service, or anything else on the HIPAA identifier list. A denial reason, a CPT code and a de-identified clinical summary are fine.

Enterprise versions with data protection terms are starting to appear, and EHR vendors will eventually deliver these features inside systems that are already covered by your BAA. Until then, treat the public tools as public. If the incident at the start of this article happened in your practice, the response is the same as for any other impermissible disclosure: document it, run the four-factor risk assessment, and decide with counsel whether it is reportable. Then fix the rule so it does not happen again.

A three-rule policy for this month

  1. No PHI in any public AI tool. Write the identifier list into the policy and train on it with examples, because "de-identified" means different things to different people.
  2. A human owns the output. Anything sent to a payer, a patient or a regulator is reviewed and signed by a named staff member who is responsible for its accuracy, exactly as if they had written it.
  3. No coding or eligibility decisions. The tools may explain, draft and summarize. They do not decide the code, the modifier or whether a service is covered.

Add a short list of approved uses so staff know what yes looks like: appeal drafts from de-identified facts, policy summaries with the source attached, patient communication drafts, spreadsheet help, and training questions. And add the policy to the annual HIPAA training, with the denial letter example, because the abstract rule did not stop the biller in the first paragraph and a concrete example might have.

A worked example, done correctly

A claim for 20610 with 99213-25 is denied CO-97, E/M bundled into the procedure. The biller writes the prompt: "Draft an appeal letter to a commercial payer for a denial of an established patient office visit billed with modifier 25 on the same day as a knee joint injection. The visit addressed new knee pain with a differential diagnosis and imaging order before the decision to inject; the note has separate sections for the evaluation and the procedure. Cite the CPT definition of modifier 25. Do not include patient identifiers; leave placeholders." The output is a solid first draft in under a minute. The biller checks the CPT language against the book, inserts the identifiers in the practice's own system, attaches the note, and signs it. That is the right division of labor.

Now the same example done wrong: the biller pastes the remittance and the visit note as they are, asks "write an appeal for this", and sends the result. Three problems follow. The tool has the patient's name and member ID. The letter cites a "CMS modifier 25 policy" that does not exist by that name. And the argument it builds is generic, because nobody told it the fact that wins the appeal, which is the separate evaluation before the decision to inject. The letter looks better than the old template. It is worse in every way that matters.

Questions we hear

Will this replace billers?

Not the ones who know payer rules. It removes some of the writing time from their day. The judgment, the follow-up calls and the knowledge of which payer does what remain human work for the foreseeable future, and the staffing shortage in billing departments is not going to be solved by a chatbot this year.

Should we buy an AI coding product?

Ask the vendor three things: what its accuracy is on your specialty measured against certified coder review, whether it has a BAA, and what happens when the code it suggests is wrong. If it is presented as a suggestion tool that a coder reviews, it may save time. If it is presented as autonomous, we would wait.

Does Revelrex use these tools?

We use language tools for drafting where no PHI is involved, under exactly the rules above, and every claim and appeal is the work of a named person on our team. Our training EHR and RCM courses teach the coding and appeal reasoning that the tools cannot supply, and the website team is happy to talk about where patient-facing automation makes sense for a practice.

What to do this month

  1. Ask the billing team, without blame, who is already using these tools and for what. You need to know before you write the policy.
  2. Write the three-rule policy with the HIPAA identifier list attached and a short list of approved uses, and have every staff member sign it.
  3. Run a fifteen-minute training using the appeal letter example done right and done wrong.
  4. If PHI has already gone into a public tool, document it and run the breach risk assessment with counsel.
  5. Put a review date in six months on the policy. The tools, and the enterprise terms available for them, are changing fast enough that a 2023 policy will need a 2024 revision.