Fort Worth 24

collapse
Home / Daily News Analysis / AI scribes are getting drug names and diagnoses wrong in NHS records

AI scribes are getting drug names and diagnoses wrong in NHS records

Sep 01, 2026  Twila Rosenbaum  4 views
AI scribes are getting drug names and diagnoses wrong in NHS records

AI scribes have arrived in NHS consultation rooms faster than the safeguards designed to protect patients. These digital assistants listen to doctor-patient conversations, draft clinical notes, and generate letters that are then placed into medical records. But a warning from Healthwatch England, the statutory patient watchdog, has revealed a pattern of errors that are easy to overlook and potentially dangerous.

A diagnosis changed by one missing word

One striking example involves a patient in England who was told they had demyelination, the nerve damage that underlies conditions such as multiple sclerosis. The test result had actually read “null demyelination,” meaning no demyelination was present. An AI scribe had dropped the word “null,” reversing the meaning of the result. The patient was left with a record saying they had a serious neurological condition when the test had been normal. This case is among those collected by Healthwatch England and highlights how small transcription differences can produce major clinical consequences.

What the watchdog found

Healthwatch England reported that 27 different AI scribes are now in use across the health service in England. These tools sit in the consulting room, listen to the conversation, and produce the note that goes into the patient record and the letter that goes to the patient. Although the marketing suggests these systems are accurate and safe, the watchdog gathered multiple examples of errors. In one case, a scribe swapped a prescribed drug for a different medicine with a similar name. This is exactly the kind of confusion that drug naming and packaging standards are designed to prevent, because similar-sounding medication names are a known risk in healthcare. In another case, a summary omitted a consultant’s instruction that the patient should seek a repeat prescription for migraine medication. Without that instruction on the letter, the patient could have been left without a prescription and no clear path to get one. A third example recorded a doctor telling a patient to continue taking Prozac, when the doctor had neither prescribed it nor discussed it during the consultation.

Why fluent notes are dangerous

The common thread across these errors is that the output looked entirely plausible. AI scribes do not typically produce broken English or strange formatting. They generate fluent, confident, and structured notes that match the style of clinical documentation. That fluency is precisely what makes the errors dangerous. A busy clinician may glance at a note and sign it off without reading every word, especially when the note looks no different from one produced by a human. The watchdog warned that inaccuracies may persist in patient records if the patient does not catch them. That places the last line of defence on the person least equipped to know what the note should have said. The patient might not know the difference between two similar drug names or might not realise that a missing phrase changes the clinical picture.

Concerns from health leaders and academics

Rachel Power, chief executive of the Patients Association, has been among those raising concerns. Her voice is joined by clinicians, including London GP Shier Ziser Dawood, and researchers such as Charlotte Blease of Uppsala University in Sweden. The objection is not to the technology itself. Many clinicians see the potential value of AI scribes in reducing the administrative burden of writing notes. The problem is the way the tools have been introduced without a safety net. Patients are not routinely told that an AI scribe is being used, and there is little opportunity for them to verify what was recorded. The concern is not that AI will replace human judgment overnight, but that the ordinary, everyday process of recording a conversation has become another place where errors can enter the system and remain unseen.

What is an AI scribe?

AI scribes, sometimes called ambient documentation assistants, combine speech recognition with language models to convert a conversation into structured text. They are intended to capture the clinical conversation and draft a note that the clinician can edit, approve, and sign. The sales pitch is that a doctor can spend more time looking at the patient and less time looking at a screen. In an NHS under intense pressure, the promise of saving an hour of typing each day has led to rapid adoption. However, the same features that make the tools useful also create risk. The AI does not simply record every word; it summarises, selects, and organises. Summarising means deciding what is important and what can be left out. When that process goes wrong, the record can be incomplete or misleading. The output may still read well, because the model has learned to produce natural-sounding text, but that does not guarantee the text is accurate.

A regulatory debate

There is currently no England-wide oversight of AI scribes. The Medicines and Healthcare products Regulatory Agency, or MHRA, has not classified them as medical devices. This leaves them outside the regulatory regime that would test them for safety and effectiveness before deployment. The MHRA published guidance in August clarifying where the line sits. A system that only transcribes what was said is not a device; a system that suggests a diagnosis or a treatment may well be. That distinction puts a great deal of weight on how each product is described by its vendor. The incentive that follows is obvious: a scribe marketed as a passive transcriber avoids regulatory scrutiny, while a scribe marketed as a clinical assistant would have to go through a longer and more expensive approval process. The same software could be described either way, and the current guidance makes that description the deciding factor.

A different kind of AI failure

Transcription error is a different failure from the kind most AI safety work anticipates. In this case, nobody was misled by a hallucinated fact or a fabricated event. A real sentence was rendered slightly wrong. But slightly wrong is sufficient when the sentence names a drug or describes a diagnosis. The problem is subtle. A record that says a patient was told to continue a medication is different from a record that says the doctor mentioned that medication. A record that says a test showed demyelination is different from one that says no demyelination was found. The error becomes permanent evidence of something that did not happen, and it can influence future care. The patient may be told the wrong thing by another clinician, prescribed a drug they should not take, or referred for tests they do not need.

Why adoption has outpaced systems

None of this fully explains why the tools spread so quickly. Clinical documentation is the administrative burden doctors complain about most. A system that reliably removes an hour of typing a day will be adopted whether or not anyone has assessed it. The pressure to see more patients, manage more messages, and keep records up to date is enormous, and AI scribes offer a practical way to ease that pressure. Some clinicians report that the tools work remarkably well and that the convenience is transformative. Others have found errors, sometimes after a patient spots them. The deeper problem is the gap between the promise and the permanent consequence. A tool adopted for its speed, unassessed because of how it is categorised, producing a document that becomes the permanent clinical record, is a chain in which no single link is obviously anyone’s responsibility.

What should happen next

Healthwatch England has pointed to a remedy that is unglamorous and probably right. Patients should be told when an AI scribe is being used and given their notes to check. That would turn an accidental safety mechanism into a deliberate one. If patients are encouraged to read their notes and report anything that looks wrong, errors that currently survive the busy glance of a clinician could be caught much earlier. Some experts argue that AI scribes should be classified as medical devices when they are used in ways that influence clinical decisions. Others say that vendors should be required to publish information about accuracy rates, error types, and the intended use of the product. Clear accountability is also needed. If a note produced by an AI scribe causes harm, it should be possible to establish who is responsible and how the error was allowed to happen. Until such measures are in place, patients in England are relying on their own vigilance to detect mistakes in records that are supposed to describe their own care. Twenty-seven products, no device classification, and a check performed by whoever happens to read their letter carefully: that is the current arrangement in the NHS in England.


Source: TNW | Artificial-intelligence News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy