
TL;DR: Whisper, OpenAI’s supposedly “advanced” AI transcription tool, is being used in hospitals to transcribe doctor-patient conversations – and it’s a mess. This thing doesn’t just mishear words, it flat-out fabricates sentences that were never spoken. Researchers have found these hallucinations at alarming rates. One University of Michigan researcher found them in eight out of every ten transcriptions. Even with clear, high-quality audio, Whisper likes to go off-script sometimes. And this is what hospitals are relying on to create medical records?
Imagine finding out your doctor’s notes included a fictional diagnosis because their AI felt like improvising.
OpenAI’s Whisper transcription software, hailed for its “near human-level accuracy,” has been caught fabricating sentences, inventing treatments, and throwing in a bit of racial commentary for good measure.
And we’re not talking about some harmless, quirky glitches here. One transcription that was supposed to be about a boy taking an umbrella turned into a story about knives and imaginary killings.
In another case, Whisper added racial descriptors where none were given, and in yet another, it introduced a fictional medication called “hyperactivated antibiotics.” The kind of hallucinations you’d expect in a bad sci-fi movie, not in a hospital.
I’m not hallucinating it, it’s all in an Associated Press story.
Hospitals are playing a dangerous game
OpenAI warns that the tool shouldn’t be used in “high-risk domains”, but apparently this doesn’t mean much to hospitals that use Whisper to transcribe doctor-patient conversations, as if these errors are just amusing side notes. So now, hospitals are letting an experimental tool that occasionally makes up facts and diagnoses write patients’ medical records. Nabla, a company pushing Whisper-based tech in the healthcare sector, has already brought this tool to over 30,000 clinicians and 40 health systems. These folks assured AP that they’re “addressing the problem,” even while Whisper deletes the original audio for “data safety.” That means no going back to check what was actually said.
Nabla’s CEO Alex LeBrun recently wrote that their tool has “many improvements specifically developed to suppress hallucinations” and that’s the crux of the problem: hallucinations can’t be completely eradicated, only mitigated. Yet here we are, pretending that some fine-tuning will be enough to safeguard patient health.
It would be one thing if Whisper hallucinated once in a blue moon. But hallucinations pop up even in high-quality recordings: a pause, a bit of background noise, and Whisper is suddenly rewriting reality.
Whisper’s hallucinations: creative writing, not reliable transcriptions
The full extent of Whisper’s hallucination problem is tough to pin down, but let’s just say it’s not a minor quirk. A University of Michigan researcher studying public meetings reported hallucinations in up to 80% of transcripts. Eight out of ten. It’s almost like Whisper is auditioning for a creative writing gig.
Another machine learning engineer reported finding hallucinations in about half of the over 100 hours he analyzed, while yet another developer encountered fabrications in nearly every one of the 26,000 transcripts he produced.
And don’t think it’s just the messy or muffled recordings throwing Whisper off its game. Even short, crystal-clear audio snippets aren’t safe from these fever dreams. One study counted 187 hallucinations in over 13,000 well-recorded samples. Whisper goes off-script even when it has no excuse to do so.
Rushing AI into healthcare: haven’t we learned?
Look, I get it. AI is the shiny new toy, and everyone wants to be the first to slap “AI-powered” on their marketing. But let’s have some common sense here. You don’t toss a half-baked AI model into mission-critical systems and hope it works out. If the tech industry wants to tinker with Whisper, go for it: in controlled, low-stakes settings. But hospitals? Really?
We’ve seen this story before: tech companies sell the sizzle before the steak’s even cooked. Face recognition, predictive policing – remember those? Both were pushed into high-stakes environments without so much as a second thought about reliability, let alone ethical implications. Now here we are again, setting AI loose in places where it can do real harm because everyone’s chasing the next “revolution.” The tech execs get their innovation badges, but who’s left to pick up the pieces when hallucinated records make their way into patient care?
This is the type of behavior that shows we’re rushing headlong into AI integration without a moment’s thought for safety. I mean, what’s the plan here? Wait for a big public scandal when AI misdiagnoses a patient? When AI tools start making things up – especially in places like hospitals – perhaps it’s time to call a timeout.
If we’re going to bring AI into sensitive fields, it needs to be vetted thoroughly. Give it a workout in less critical settings first. Test it, fix it, test it again, and make sure it can actually handle the job before tossing it into healthcare. This idea of pushing untested, experimental AI into hospitals is reckless. And for what? A few minutes saved on transcription? In medicine, we can’t afford hallucinated “efficiency.”
You’re welcome to keep testing in your lab, but don’t bring it into the ER.