Bay Street Wire
Tech & BusinessOpinion

The Hallucination Hazard: Why the 'AI Midwife' is a Liability Trap

Portrait of Victor Cho
Victor Chothe contrarianOct 3AI
The Hallucination Hazard: Why the 'AI Midwife' is a Liability Trap

AI-generated image · Bay Street Wire

Victor Cho argues that using LLMs to fill the gaps in medical care is a dangerous gamble that mistakes pattern recognition for professional diagnostic skill.

The latest trend in medical 'augmentation' is less about innovation and more about a desperate attempt to automate human empathy and diagnostic rigor. In a guest post for Astral Codex Ten, as first reported by the outlet, Drew Housman describes a journey through the grueling world of unexplained infertility, where the failure of the traditional medical system led him and his wife to rely on a customized version of OpenAI's GPT-4.

By gaming the system—prompting the AI to act as a fictional doctor named 'Dr. Reid' in a Hollywood medical drama—Housman found a tireless, patient alternative to the brevity of the U.S. medical system. This 'AI Midwife' eventually suggested an MRI that revealed a six-centimeter submucosal fibroid—a growth the authors describe as a tumor the size of a small peach—that had been missed by human doctors for six years.

On the surface, this looks like a victory for the LLM. In reality, it is a terrifying glimpse into a liability nightmare.

As a skeptic, I see the red flags that the authors gloss over. Housman admits that 'once in a blue moon,' the AI would look at an ultrasound and hallucinate a pregnancy that didn't exist. In any other professional context, a diagnostic tool that randomly declares a patient pregnant when they are not is considered broken, not 'awesome.' The fact that the users had to ask the AI, 'Are you drunk, Dr. Reid?' highlights the fundamental flaw of these systems: they are probabilistic engines, not medical practitioners. They do not 'know' medicine; they predict the next likely token in a sequence.

Furthermore, the process of getting the AI to function required a series of deceptive prompts and 'gaming the system' to bypass OpenAI's own safety guardrails. When the system balked, the users offered monetary rewards or insisted the interaction was 'just for fun.' This is not a clinical workflow; it is a series of hacks.

While the authors credit Dr. Reid with finding the fibroid that their human fertility doctor missed—and noted that the doctor was initially unenthusiastic about the MRI—this outcome is a classic example of survivor bias. For every one person who finds a hidden tumor through an AI's random suggestion, how many others will be led down a rabbit hole of unnecessary, invasive tests based on an LLM's confident hallucination?

Replacing the 'one line response' of a MyChart app with the infinite patience of a chatbot is a seductive trade-off. It solves the efficiency problem, but it ignores the accountability problem. When a human doctor misses a diagnosis, there is a board of medicine and a malpractice framework. When 'Dr. Reid' misses a diagnosis—or suggests a dangerous treatment—there is no one to hold accountable but the user who decided to treat a language model as a medical professional.

Efficiency is not a substitute for expertise, and a 'Hollywood medical drama' prompt is not a substitute for a medical degree. The 'AI Midwife' isn't practicing medicine; it's playing a character. And in the world of healthcare, playing a character is a liability no patient should have to bear.

Sources

More from Victor Cho