The Automation Paradox: Why AI Struggles Where It Should Excel
In this episode of What is Your Problem?, radiologist Saurabh Jha explains why AI has failed to change radiology despite a decade of hype. The core issue is that we are optimizing for the wrong variable. While developers focus on pattern matching, they ignore the reality of clinical practice, where the hardest task is not identifying disease, but correctly identifying what is normal. This conversation shows how immediate, low-stakes AI deployment creates automation bias that degrades human expertise. For leaders and practitioners, the lesson is clear: AI is not a plug-and-play productivity tool. It is a high-stakes cognitive investment that requires deep expertise to govern. Those who use AI to bypass foundational learning will become less capable, while experts who use it to augment their own judgment will build a lasting competitive advantage.
The Hidden Cost of Helpful Flags
The idea that AI would replace radiologists was built on the assumption that radiology is just pattern matching. But as Jha explains, the system responds to AI with increased cognitive load rather than efficiency. When AI flags potential issues, it forces the radiologist to perform two tasks: assess the patient and assess the AI output.
For most parts it was not really finding things that I was not finding but it was finding things that were not real. So it increased my cognitive burden.
-- Saurabh Jha
This creates a feedback loop of false positives that mirrors the boy who cried wolf. In the past, computer-aided detection systems in mammography led to more biopsies and patient anxiety without improving health outcomes. The system failed because it optimized for sensitivity at the expense of specificity. The result is that practitioners stop trusting the tool, making the technology an expensive, irritating distraction rather than a force multiplier.
The Expertise Trap: Why Novices Are at Risk
Systems thinking requires us to look at how a tool changes the user over time. Jha highlights a dangerous dynamic for trainees: automation bias. If a trainee relies on AI before they have developed the internal, intuitive expertise to distinguish disease from the fuzzy border of normal, they never actually learn the craft.
If you are starting off, you are nervous... there is a big chance AI can make you worse and that is because of a phenomenon called automation bias where if you do not know enough... you start relinquishing more and more things to AI because you never had the point where there was no way where you were making those decisions on yourself.
-- Saurabh Jha
When the drudgery of entry-level work is automated, the scaffolding of expertise is removed. The long-term cost of this shortcut is a generation of practitioners who cannot function without the machine, creating a brittle system that collapses the moment the AI makes a mistake.
Where Immediate Pain Creates Lasting Moats
The most successful applications of AI are not found in resource-rich hospitals, but in settings where no alternative exists. Jha’s experience at Everest Basecamp and in rural India shows that AI becomes a transformative technology only when it solves a total absence of expertise.
In these environments, AI acts as a median doctor, providing a baseline of care that is better than the alternative. The system succeeds here because the trade-off is clear: the gain of AI intervention outweighs the risk of false positives. This suggests that AI’s greatest value is not in disrupting high-functioning experts, but in filling systemic voids where the cost of doing nothing is higher than the risk of an imperfect machine.
Key Action Items
- Audit your drudgery: Identify which tasks are foundational for learning versus which are administrative. If you automate the foundational tasks, you must replace them with rigorous, AI-free testing to ensure expertise is still being built. (Immediate)
- Shift from accuracy to utility: Stop asking if your AI model is accurate in a vacuum. Ask if it reduces the total cognitive load of the operator. If it creates more flags, it is failing, regardless of its technical precision. (Next Quarter)
- Implement human-in-the-loop constraints: For high-stakes decisions, mandate that the human must form a conclusion before consulting the AI. This prevents the brain from defaulting to the machine suggestion. (Immediate)
- Build for the normal: If you are developing or deploying diagnostic tools, prioritize the ability to identify normal variants. Over-calling normal is the primary cause of system failure in diagnostic fields. (12-18 months)
- Audit for automation bias: If you are a leader, test your team’s ability to perform their core functions without the AI. If they cannot, your system is creating a vulnerability, not an advantage. (Next 6 months)