VA let general-purpose AI chatbots into clinical work without high-impact safeguards

Tombstone icon

A Department of Veterans Affairs inspector general review found that staff were using VA GPT and Microsoft 365 Copilot Chat for clinical notes and patient-care summaries even though VA had not classified either tool as high-impact AI. That decision spared the chatbots from safeguards applied to a separate AI scribe, including pre-deployment testing, impact assessments, ongoing monitoring, and additional human oversight. VA also lacked reliable ways to label AI-generated records or report AI-related safety events. The watchdog documented a production clinical-governance hazard, but no confirmed patient injury.

Incident Details

Severity:Facepalm
Company:Department of Veterans Affairs
Perpetrator:Public agency leadership
Incident Date:
Blast Radius:Clinical notes and care summaries created with two VA-wide chat tools; actual clinical usage and patient impact could not be measured under existing monitoring

General-Purpose Tools, Clinical-Purpose Work

On June 11, 2026, the Department of Veterans Affairs Office of Inspector General published the results of a national review into generative AI chat tools used by the Veterans Health Administration. The review examined work performed from October 2025 through early February 2026 and followed a preliminary warning issued in January. Its conclusion was unusually plain for an oversight report: VA had provided clinical staff with generative AI chat tools for clinical use without the oversight or safeguards needed to monitor the patient-safety risks.

The two tools were VA GPT, built for general use inside the department, and Microsoft 365 Copilot Chat, also available broadly to VA staff. Neither was designed as a clinical product. Staff could nevertheless enter patient information, ask the tools to draft notes or summarize care, and use the generated text in clinical documentation. The OIG found that VA supported this work with presentations, training materials, prompt-writing guidance, Microsoft Teams communities, and a staff-created prompt-sharing application.

This was production use rather than a laboratory exercise. Its exact scale was unknowable because VA had no central measure of clinical use. Activity in two AI-focused Teams sites offered only a rough indicator: one site had 10,997 active users over a 90-day period, while another had 4,835. Those figures describe community activity, not a count of clinicians placing chatbot output in medical records. In a sample of 135 voluntarily shared prompts, however, investigators classified 79 as clinical. Fifty-six concerned clinical notes, 17 summarized information, and six served other clinical purposes.

High Impact, Unless It Arrived Through a Chat Window

Federal guidance defines high-impact AI partly by whether it forms a principal basis for decisions or actions affecting human health and safety. Patient diagnosis, risk assessment, and treatment are named examples. Classifying a use that way triggers controls such as pre-deployment testing, an impact assessment, ongoing performance monitoring, staff training, additional human oversight, and mechanisms for accountability.

VA had applied that classification to Ambient AI Scribe, a targeted tool that listens during patient visits and drafts medical-record notes. It had not classified VA GPT or Copilot Chat as high impact, even when staff used them to perform similar documentation work. The interface apparently carried more regulatory weight than the job being done. A purpose-built scribe received the serious controls; a general chatbot drafting comparable text did not.

VA AI leaders compared the chat tools to search engines. The OIG rejected that analogy. A search engine returns links for a person to inspect. A generative model synthesizes and transforms source material into new prose, sometimes with invented details or omitted facts presented in the same confident voice as accurate material. Calling that process a search does not make it one. It merely gives a risky workflow a familiar name and hopes the clinical consequences respect the branding.

The OIG cited medical-documentation research showing why the distinction matters. One study reported a 1.5 percent hallucination rate; 44 percent of those hallucinations were judged major enough that they could affect diagnosis or patient management if left uncorrected. It also reported a 3.5 percent omission rate for sentences containing relevant clinical information, with 17 percent of those omissions considered major. Those were research results, not measured error rates for VA GPT or Copilot Chat. VA had not established the testing needed to produce equivalent figures for its own clinical workflows.

A Safety System That Could Not See AI

The governance gap continued after a note was created. VA's National AI Institute and Chief AI Officer team had deployed and supported the chat tools with little coordination from the Veterans Health Administration's National Center for Patient Safety. The patient-safety center's AI lead told investigators that the office should have been involved even if only for awareness, but attempts to establish a joint working arrangement had gone nowhere before the review.

VA uses the Joint Patient Safety Reporting system to collect medical errors, close calls, and near misses. According to the OIG report, it receives about 180,000 submissions each year. In 2025, the patient-safety center asked the Defense Health Agency, which shares the platform, to add AI as a selectable cause. No formal decision had been reached when investigators reviewed the system.

The OIG searched those reports for mentions of AI and found none. That finding does not establish that AI caused zero incidents. A safety manager could categorize an event by the patient's outcome without recording that a chatbot contributed, and VA had no labeling process to identify AI-generated documentation later. A reporting system with no AI category is very efficient at producing no AI-category incidents. It is considerably less useful for finding patterns before they become injuries.

The same problem affected the source material. VA did not centrally curate or evaluate clinical prompts or the outputs they produced. General tips existed, but investigators found no organized process for refining prompts, measuring their error rates, or identifying unsafe prompt patterns across facilities. Each clinician remained responsible for reviewing output, while the institution lacked the feedback loop needed to learn whether the same failure was recurring elsewhere.

What the Report Did and Did Not Prove

Neither the final OIG report nor Nextgov's independent coverage identifies a veteran harmed by an AI-generated note or recommendation. The report documents a production clinical-governance hazard, not a confirmed patient-injury case. That boundary matters. A broken reporting process cannot be converted into evidence of injuries, but it also cannot support reassuring claims that none occurred.

After the January preliminary memorandum, a VA official told Nextgov that clinicians used AI only as a support tool and that appropriate VA staff always made patient-care decisions. Human review is an important safeguard, and the OIG did not claim the chatbots were independently diagnosing or treating veterans. The concern was that VA had placed nearly all responsibility at the user level while skipping institutional controls applied to another tool doing closely related work.

Human review also needs a system around it. Clinicians need clear permissible-use rules, tested prompts, training tied to specific workflows, visible labeling, quality checks, and a way to report when generated text is wrong. Without those pieces, "a human checks it" becomes a policy that cannot be measured. Every error depends on one busy person spotting it before generated prose settles into the medical record and begins influencing later care.

VA's Response

The Under Secretary for Health concurred in principle with the OIG's first recommendation and concurred with the other two. VA agreed to define permissible clinical uses, assign oversight responsibilities, examine whether safeguards used for high-impact systems should apply to the chat tools, and integrate AI risk into patient-safety monitoring and staff training.

VA's action plan calls for pre-deployment performance assessments, impact assessments, ongoing monitoring, use-specific education, transparency mechanisms, and quality-assurance processes. It also says the National Center for Patient Safety is working to update reporting systems so AI risks and incidents can be tracked. The target completion date for all three recommendations is April 2027.

Some repairs had already begun by the final report. Patient-safety representatives were added to an AI assessment subcommittee, VA started coordinating with the Defense Health Agency on an AI-specific reporting category, and officials planned education on recognizing AI as a contributing factor during root-cause analysis. The OIG said it would follow up to determine whether those measures were effective and sustained.

The gap exposed here came from classifying risk according to the product label rather than the clinical function. VA recognized that an automated scribe drafting a medical note deserved high-impact safeguards. It treated a general-purpose chatbot drafting a medical note more like a search box, despite the output landing in the same consequential workflow. Hospitals do not become safer when an untested clinical function wears office-software clothing. They merely become worse at seeing where the risk entered.

Discussion