Key takeaways
- Clinical AI needs rigorous validation before deployment and continuous monitoring afterward, because performance can change as patient populations and clinical environments shift.
- A model that performs well in one population may fail in another, making local validation and representative data essential.
- Governance must include real authority to monitor, retrain, or stop a model when it no longer performs safely.
- Nurses and patients should be part of AI oversight, because frontline clinical judgment and patient trust are critical to safe deployment.
- AI should be built and nurtured locally, with institutions investing in the teams and systems needed to train, test, validate, and supervise models over time.
At the Community Oncology Global Congress (COGC 2026), organized by OncoDaily, Iyad Sultan, Chief Health Informatics Officer at King Hussein Cancer Center (KHCC) and pediatric oncologist, focused on one of the most difficult questions surrounding artificial intelligence in healthcare: how do we know that an AI model that performs well today will remain safe when it meets real patients tomorrow?
At KHCC, that question is already practical rather than theoretical. The institution has built its own AI office, data infrastructure, extraction pipelines, dashboards, and early clinical decision-support systems. But building the model challenges to be only the beginning. The harder work is validating it, monitoring it as populations change, and creating governance capable of stopping it when it no longer performs as expected.
Building AI From the Ground Up
“We built an AI office in a part of the world where very few of these exist. We built our data infrastructure, created a mirror of our electronic medical record, and developed extraction pipelines.
We now have more than 100 pipelines running every night, feeding dashboards and sending daily notifications. We have also started developing our first decision-support system, which we call the Intelligent Patient Navigator.
Our models are being developed with academic partners, engineers, people within the institution, and champions across different departments.
We are already using agents during development, and eventually we expect agents to become patient-facing as well – potentially helping answer questions, file claims, and support other tasks.
But healthcare AI development has to follow a much more rigorous process than technology in many other industries. In another business, a poorly performing system may lose money. In healthcare, a life can be lost.”
Healthcare AI Needs a Full Development Lifecycle
“The development process needs structure.
Version control allows us to return to earlier versions of a model or prompt and understand how a change affected the outcome.
Then there is training and fine-tuning. Fine-tuning can often be done with relatively small amounts of data and can substantially change how a model behaves for a particular clinical problem.
Generative AI creates another challenge because, unlike deterministic models, the same question can produce different answers. That is one of the remarkable properties of these systems, but in medicine it also creates obvious validation problems.
We should also not forget simpler tools. Encoder models can act almost like watchdogs trained for a very specific task. For example, a model can monitor the record and flag a possible adverse drug reaction without needing to behave like a general-purpose language model.
And all of these systems depend on data. Synthetic and semi-synthetic data can help when real patient data are limited, but nothing truly replaces the quality of real patient data and real-world experience.”
A Good Model Can Fail in a New Population
“One of the clearest examples of how things can go wrong comes from a sepsis prediction model developed for emergency department patients.
The model initially reported an area under the curve of approximately 0.76, which looked quite good.
But when investigators evaluated it across more than 30,000 emergency department visits at the University of Michigan, performance fell to approximately 0.63.
The model missed about 67% of patients who actually developed sepsis. It generated alerts in approximately 18% of emergency visits, and only about one in eight of the patients triggering an alert actually had sepsis.

So a model that appeared to perform well under one set of conditions performed very differently when it encountered another population.
Why does that happen? One major reason is the data.
The population used to train the model may not resemble the population where the model is eventually deployed. Many open datasets contain substantial demographic and socioeconomic biases.
We can also choose the wrong endpoint. If we use healthcare spending as a proxy for disease severity, for example, we may actually be measuring differences in access and spending rather than differences in illness.
The model can therefore be mathematically correct and still clinically wrong.”
Validation Does Not End at Deployment
“Validation is a much bigger process than reporting an AUC, F1 score, sensitivity, or specificity.
We begin with analytical validation. Then the model has to undergo clinical validation, where it is tested against actual patient data – ideally both locally and, where appropriate, in external populations.
And even after deployment, validation has to continue.
The lifecycle of an AI model never stops.
What works today may stop working tomorrow because the population changes, clinical practice changes, or the model itself begins to drift.
Approval also does not mean that a model is useful. Validation may demonstrate that a model is accurate, but that does not automatically demonstrate that clinicians will use it or that it improves care.
Accuracy and usefulness are not the same question.”
What Safe AI Governance Requires
“Governance is really a balance.
Patients have the right to privacy, but they also have the right to access models that may improve or even save their lives.
Investigators need freedom to innovate, but patients must also be protected from harm. Physicians need to be able to use these tools safely, while at the same time we may eventually face questions about liability when a validated tool exists and a clinician chooses not to use it.
There is no shortage of published AI governance frameworks. What is often missing is something much more basic: who actually has the authority to govern the model?
Different regions are approaching the problem differently.
In the United States, the FDA often approaches clinical AI through existing medical-device pathways. The European Union is moving toward greater mandated transparency under the EU AI Act. South Korea has developed a more direct government supervisory role.

There is an enormous opportunity for low- and middle-income countries.
We can build systems capable of working with imperfect medical documentation and supporting clinicians where specialist resources are limited. There are already examples showing that AI can improve care.
But simply importing regulatory frameworks from the United States or Europe is not necessarily the answer.
These technologies are moving extremely quickly, and we need governance structures that recognize both their risks and their potential to save lives.
Language models are particularly challenging because they behave differently from conventional deterministic software. They can make mistakes, and the context in which they are used can sometimes produce more problems than the model itself.
The important question is therefore not only whether the model works technically, but whether it works inside the clinical pathway where we intend to use it.”
Model Drift: Keeping Humans Involved
“We built an emergency department model at KHCC. We extracted the data, developed the model, validated it, and we are now moving through deployment and post-deployment supervision.
But we already know that this model will eventually experience the same problem every model faces: drift.
New regulations this year mean that KHCC will begin seeing more older patients. That means the population using our hospital will change.
A model trained on the previous population may therefore begin performing differently, and it may need to be retrained.
That is a very simple example of why these systems cannot simply be deployed and left alone.
The patients change. Clinical practice changes. The environment changes. The model has to change with them.
Any model that does not have nurses involved in its supervision may ultimately do more harm than good.
Nurses at the bedside are often the people best positioned to recognize when a model is drifting or when its recommendations no longer match what is happening clinically.
Any serious hospital AI governance structure should therefore have nursing representation as a central part of the system.
Patients also need to be involved.
Survey data from more than 3,000 patients showed greater trust when people knew that an AI model had been tested and performed at or above specialist level. Trust was also higher when a clinician remained involved, when the training data represented people like them, and when they knew there was real institutional governance overseeing the system.
And governance is not simply having a committee.
It means having a body that can monitor a model, understand when it is failing, and stop its use when necessary.”
Local AI, Local Governance
“My final piece of advice is that AI cannot simply be bought. It has to be nurtured locally.
Institutions need an office, committee, or other dedicated structure that is actively training, testing, validating, and monitoring models before and after deployment.
Governance also cannot be universal. It will differ across countries and health systems, and it will continue changing over time.
This is a moving target. Models will improve, hallucinations may become less frequent, and accuracy will continue to increase.
But our goal has not changed.
We are still trying to save more lives and deliver better medicine.
We do not know exactly what comes next, but we should use these models thoughtfully, supervise them carefully, and make sure that technological progress ultimately translates into better care.”
Written by Eliz Baloyan, MD, Features Writer and Editor at OncoDaily and CancerWorld