Allen Li, Medical Oncologist at The Vancouver Clinic and Creator of Oncology AI Lab, shared on Substack:
“Every few months another study crowns the best medical AI, and clinicians line up behind their favorite like a team they drafted. Which one wins is not the question you actually have. The contest these studies run, which chatbot scores higher on a set of questions, cannot tell you the thing you need to know at the bedside: whether to trust this answer, for this patient. An AI comparison can rank answers. It cannot make that call. It was always yours.
Vinay Prasad, a hematologist-oncologist, showed what that judgment looks like by reading the actual queries in one of these studies. One asked how to adjust the chemotherapy for a patient with metastatic cholangiocarcinoma who was doing poorly. The sharpest move there is not to pick a dose. It is to ask whether more chemotherapy is the right goal at all, and whether this patient is better served by a conversation about what comes next. That is not giving up on the patient. It is care. No ranking of answers can reach it, because the skill is not choosing the better answer. It is seeing that the question is wrong. A preferred answer is not a correct one, and for many questions in medicine there is no single correct one at all.
Two of these studies landed this year, and they disagreed. One, in Nature Medicine, found the general models beating OpenEvidence and UpToDate. The other, a preprint built on OpenEvidence’s own platform, found the reverse. These are serious studies, and independent evaluation of tools already entering practice is worth doing. A benchmark is how you catch the tool that invents a citation or misses a red flag before it reaches a patient, and that safety floor matters. The disagreement between them is not random either: each study’s winner was the one whose home turf it ran on, which is a story worth its own telling. But set all of that aside, because it is not the deepest problem. Suppose the studies had agreed. Suppose there were one clean, incorruptible leaderboard, and it named a winner. It still would not tell you whether to trust the answer on your screen, for the patient in front of you. That question is not on the test.

The judgment is yours
For a practicing clinician, the tool is a reference, not a verdict. It is a faster way to remember what you half-remember, or to surface an option you had not considered. Whether that output is right, for this patient, in this room, is your call.
Which is why lining up behind a favorite AI gets it backwards. We do not root for the control or experimental arms in a clinical trial. We ask what the design shows, not which side we are on. A tool is not a jersey, and a leaderboard is not a loyalty test. Choosing a chatbot to defend is a way of handing your judgment to a scoreboard, and that is the one thing you should never do with it.
I want AI to succeed in medicine, and it will not get there on a leaderboard. It gets there when a tool is good enough to earn the trust of someone who can tell the difference. That someone is you, the clinician. You already know a strong trial from a weak one. You already know when a literature search has found the right paper and when it has missed. You have been deciding what to trust, for the patient in front of you, your whole career. Reading an AI output is that same skill, not a new one. A benchmark can tell you which output other clinicians preferred. Whether the one in front of you is right, you can see for yourself.
And part of that same training is knowing when you do not know. Recognizing that a question belongs to another specialty, that a case is past what you should settle alone, that a patient needs a second opinion, is not a gap in your skill. It is one of your skills. You were taught it in training, you have used it ever since, and you call it every time you order a consult. And no AI, and no benchmark, can do that for you. Reading an output and knowing when to reach past it are both yours, and both are the job.
When the next benchmark race crowns the next winner, you can let the result pass. The leaderboard cannot answer the question in the room: is this output right for this patient.”

Other articles about AI in Oncology on OncoDaily.