Arturo LoAIza-Bonilla, Network Chief of Hematology and Oncology at St. Luke’s University Health Network, and Co-Founder and Chief Medical AI Officer (CMAIO) of Massive Bio, shared a post on LinkedIn:
“The AI risk that reaches your patients first won’t look like a catastrophe. It will look like a slightly worse triage decision – repeated ten thousand times.
This week Demis Hassabis published a serious framework for governing frontier AI: a rigorous, adaptive Standards Body, pre-deployment testing, and ‘cautious optimism’ as the right posture when the stakes are high.
I agree with the instinct. But most patients won’t meet AGI first. They’ll meet the narrow, predictive, and consumer AI already sitting in triage, imaging, documentation, and decision support today.
That’s the gap Nikhil Thaker and I set out to close in a new Perspective in Artificial Intelligence in Health: From P(doom) to P(harm).
P(doom) is the probability of catastrophic, civilization-scale AI failure.
P(harm) is its nearer, testable companio: the probability that an AI-exposed clinical encounter leads to harm. In medicine, that harm rarely arrives as one dramatic event. It accumulates, through small, repeated shifts in diagnosis, clinician behavior, training, and workflow:
- AI-clinician discordance in the most complex patients (72.5% concordance with the tumor board, worse in the sickest)
- Automation bias: GPT-4 alone beat physicians on diagnostic reasoning, yet physicians using it did no better than those without it
- Skill erosion: colonoscopy detection rates dropped in unassisted cases after clinicians grew used to the AI
- Consumer triage failure: one structured evaluation under-triaged 52% of true emergencies
- Correlated, systemic failure when thousands of institutions quietly share one model.
You cannot catch drift like this with a pre-release benchmark or a one-time model card. It only surfaces in the live workflow, after deployment. So we propose P(harm) as an operational surveillance layer (workflow-indexed, severity-weighted, independently adjudicated) paired with a clinical AI safety committee that holds real authority to pause, revalidate, roll back, or retrain.
Demis argues for governing AI at the frontier. This is the same instinct one altitude down, at the floor of care.
Because healthcare does not have to choose between optimism about AI and vigilance about safety. The future of AI in medicine will depend less on model performance in isolation than on whether we keep human judgment accountable, auditable, and resilient inside the human-AI dyad.
Read the paper.
Hassabis’s essay.”
Other articles featuring Arturo LoAIza-Bonilla on OncoDaily.