Fabio Ynoe de Moraes, Associate Professor at Queen’s University, shared on LinkedIn:
“1. One-click functional-lung avoidance planning
Xiong T, Chen Z, Fan C, et al. Clinical implementation and evaluation of an artificial intelligence-driven one-click automatic planning system for functional lung avoidance radiotherapy. Pract Radiat Oncol. Published September 16, 2026.
Title: Clinical implementation and evaluation of an artificial intelligence-driven one-click automatic planning system for functional lung avoidance radiotherapy
Authors: Tianyu Xiong, Zhi Chen, Chengcheng Fan, Chunyu He, Bing Li, Guangping Zeng, Jiaxing Han, Vincent W. S. Leung, Yu-Hua Huang, Zongrui Ma, Ge Ren, Yang Sheng, Qingrong Jackie Wu, Hong Ge, Jing Cai
Read the Full Article on Practical Radiation Oncology
Study type: Retrospective planning study following TPS integration.
Population and methods: AP-FLART combined automated beam-angle selection, multimodal dose prediction, and function-guided dose mimicking in RayStation. It was tested in 33 lung-cancer patients with SPECT ventilation or perfusion imaging. Automatic plans were compared with manually generated conventional and functional-avoidance plans; three clinicians performed blinded reviews.
Principal findings
- High-function-lung mean dose decreased by 15.1% versus manual conventional plans.
- In patients classified as likely to benefit, modeled grade ≥2 pneumonitis risk fell by 6.25 absolute percentage points, a 27% relative reduction.
- 87.9% of automatic plans were acceptable without modification.
- 68.7% were judged comparable (38.4%) or superior (30.3%) to manual functional-avoidance plans.
- Planning time decreased from 2–3 hours to approximately 8 minutes.
What changed: This moves functional avoidance from bespoke expert planning toward an integrated, near-push-button workflow. The combination of functional imaging, automated beam selection, and deliverable optimization is more clinically meaningful than dose prediction alone.
Why it matters: Planning complexity is an important barrier to trials and adoption of functional avoidance in locally advanced lung cancer. Automation could make prospective testing feasible across more centers.
Limitations and risk of bias: Despite ‘clinical implementation’ in the title, the abstract does not establish prospective treatment delivery or observed reductions in pneumonitis. NTCP improvement was modeled, not clinical. The 33-patient dataset is small, and the “benefiting” subgroup may be vulnerable to post hoc selection. Almost one-third of automatic plans were rated inferior to manual FLART, even if many remained acceptable.
Verdict: Must read for thoracic RT and automated-planning investigators; not yet evidence to adopt FLART routinely.
2. On-premise clinical agents with selective autonomy
Zhang L, Wölflein G, Ferber D, et al. On-premise medical AI agents for reliable clinical decision-making. Nat Med. Published September 15, 2026.
Title: On-premise medical AI agents for reliable clinical decision-making
Authors: Li Zhang, Georg Wölflein, Dyke Ferber, Junhao Liang, Zunamys I. Carrero, Xuewei Wu, Julien Vibert, Jan Clusmann, Lino Möhrmann, Elena E. Möhrmann, Catharina Wichmann, Fabian Wolf, Tim Lenz, Jakob Nikolas Kather
Read the Full Article on Nature Medicine

Study type: Clinical-agent development and benchmark evaluation.
Data and methods: A locally deployed physician agent interacted with a simulated patient agent, ordered investigations, and produced diagnoses and reasoning traces on two MIMIC-IV-derived benchmarks. Reliability was estimated using likelihood, linguistic certainty, and consistency across five stochastic runs.
Principal findings
- Diagnostic accuracy was 90.04% on a seven-disease task and 83.8% on a four-disease task.
- Behavioral consistency best discriminated correct from incorrect diagnoses: AUC 0.860.
- Performance persisted under stress testing: AUC 0.875.
- At a prespecified consistency threshold of 0.90, the system retained 49.4% of cases with 98.9% diagnostic accuracy, deferring the remainder.
What changed: The most important advance is not raw accuracy. It is the combination of local institutional control with a decision-time abstention mechanism—an architecture potentially applicable to oncology MDT and treatment-selection systems.
Why it matters: Oncology agents should not answer every case. A clinically credible system needs to recognize a lower-risk operating envelope and escalate uncertain, unstable, or out-of-distribution cases.
Limitations and risk of bias: These were simulated encounters derived from retrospective ICU data, not prospective patient care or oncology practice. Consistency is not equivalent to correctness—a model can be reproducibly wrong. Five-run stability testing adds latency and compute requirements. Threshold performance requires external validation, and no automation-bias or clinician–AI interaction was evaluated.
Verdict: Must read for LLM safety and clinical decision-support researchers; conceptually important, not evidence for autonomous oncology decisions.
3. CT prediction of tertiary lymphoid structures in pancreatic cancer
Yu H, Li X, Cheng X, et al. Deep learning-based CT model for non-invasive prediction of tertiary lymphoid structures in pancreatic cancer: a multicenter study with prospective validation in an immunochemotherapy cohort. J Immunother Cancer.
Title: Deep learning-based CT model for non-invasive prediction of tertiary lymphoid structures in pancreatic cancer: a multicenter study with prospective validation in an immunochemotherapy cohort
Authors: Haopeng Yu, Xiaoying Li, Xuan Cheng, Yan Deng, Yuqi Wang, Yu Li, Bingjie Liu, Zixing Huang, Dan Cao
Read the Full Article on J Immunother Cancer.

Study type: Multicenter biomarker development with prospective-cohort validation.
Population and methods: A ResNet50 model was developed using 223 resected PDAC patients from two centers, clinically validated in 82 nonsurgical patients, and then locked and evaluated in 44 patients enrolled in a prospective immunochemotherapy study. Multiplex spatial profiling investigated biological correlates.
Principal findings
- TLS-prediction AUC: Training: 0.987 Internal validation: 0.937 External validation: 0.853 Prospective pathology-anchored cohort: 0.929
- In the nonsurgical cohort, predicted TLS positivity was associated with median OS of 19 versus 7 months.
- In the prospective treatment cohort, DL-high patients had: Objective response: 90.0% versus 26.5% Median PFS: 13.1 versus 5.3 months Median OS: 17.6 versus 8.7 months
- The imaging score was unrelated to PD-L1 but correlated with spatial dendritic-cell–T-helper and T-helper–B-cell interactions.
What changed: Unlike many imaging biomarkers, this model combines external validation, a locked prospective application, outcome associations, and spatial biological grounding.
Why it matters: A CT-derived TLS biomarker could enable noninvasive immune-phenotyping where repeated tissue acquisition is difficult and may help enrich future PDAC immunotherapy trials.
Limitations and risk of bias: The prospective cohort contained only 44 patients. The decline from training to external AUC indicates meaningful domain sensitivity. Very large multivariable effect estimates—PFS HR 0.079 and OS HR 0.052—raise concern about sparse events, cut-point optimism, or model instability. Because there was no untreated or randomized comparator, the study establishes prognosis and treatment-response association—not prediction of immunochemotherapy benefit.
Verdict: Must read for imaging-biomarker and immunotherapy investigators; requires randomized treatment-interaction validation.
4. Transporting deep-learning proton planning to an external center
van Bruggen IG, Wolf AL, Kroesen M, et al. Applying a pretrained DL-based IMPT planning workflow in an external center: A feasibility study in oropharyngeal cancer. Radiother Oncol. 2026;:111782.
Title: Advanced dose accumulation methods: A comparative study in online adaptive radiotherapy for abdominal-pelvic lymph node oligometastases
Authors: Erik van Lieshout, Joost J.M.E. Nuyttens, Remi A. Nout, Mischa S. Hoogeman, Maaike T.W. Milder
Read the Full Article on Radiotherapy and Oncology

Study type: External-center transfer and feasibility study.
Population and methods: A pretrained U-Net predicted robustly optimized IMPT dose distributions. Mimicking parameters were adapted using five oropharyngeal cases over two working days and evaluated in ten independent patients using quantitative metrics and blinded multidisciplinary review.
Principal findings
- Manual plans were preferred in 40%, automated plans in 40%, and plans were equivalent in 20%.
- Only 6/10 plans in each group were considered clinically acceptable.
- Automated plans reduced: Oral-cavity mean dose: 31.5→30.2 Gy(RBE); p=0.002 Parotid mean dose: 20.1→17.8 Gy(RBE); p<0.001
- Modeled xerostomia NTCP decreased: Grade ≥2: 38.1%→37.1% Grade ≥3: 10.3%→9.9%
What changed: The study addresses an underexamined problem: whether a pretrained planning model can be adapted rapidly outside its originating institution.
Why it matters: Transferability is essential if automated planning is to disseminate proton-planning expertise rather than remain a single-center technology.
Limitations and risk of bias: Five adaptation cases and ten test cases cannot characterize rare or severe failures. The statistically significant NTCP differences were only 1.0 and 0.4 percentage points, with uncertain clinical value. That just 60% of both automated and manual plans were considered acceptable complicates any claim of successful deployment. Transfer occurred within a compatible planning ecosystem and does not establish multivendor portability.
Verdict: Worth scanning—important workflow concept, preliminary evidence.
5. GPT-assisted radiomics development in head-and-neck cancer
Shi H, Lan T, Gao X, et al. GPT-assisted radiomic modeling for predicting pathological complete response to neoadjuvant chemoimmunotherapy in head and neck squamous cell carcinoma. Phys Med Biol. Published September 14, 2026.
Title: GPT-assisted radiomic modeling for predicting pathological complete response to neoadjuvant chemoimmunotherapy in head and neck squamous cell carcinoma
Authors: Huaxian Shi, Tianjun Lan, Xiaoling Gao, Hanxiao Song, Ying Li, Yu Wang, Lingjie Yang, Xiaohua Ban, Xi Zhong, Chaobin Pan, Xiaohui Duan, Yu Peng, Zhaoyu Lin, Dong Zeng
Read the Full Article on Physics in Medicine & Biology

Study type: Comparative model-development study with multicenter prospective validation.
Population and methods: Pretreatment T2-weighted MRI from 186 training, 116 validation, and 269 prospective multicenter patients was used to predict pathologic complete response. GPT assisted preprocessing, feature selection, code generation, hyperparameter optimization, execution, and output generation; it was not the classifier. Each workflow was repeated five times.
Principal findings
- For fused radiomic/deep-learning features, manual versus GPT-assisted logistic-regression AUCs were: Training: 0.759 versus 0.763±0.003 Validation: 0.714 versus 0.741±0.008 Prospective cohort: 0.700 versus 0.706±0.004
- Results were broadly similar across five classifier families.
- Repeated GPT-assisted runs showed AUC SDs of 0.001–0.019.
What changed: This evaluates an LLM as a research-workflow agent rather than as the diagnostic model. It suggests that supervised GPT workflows can reproduce conventional radiomics pipelines without sacrificing performance.
Why it matters: Research automation could reduce coding barriers and accelerate multicenter modeling, particularly in settings without extensive data-science support.
Limitations and risk of bias: The resulting clinical model remained only moderately discriminative externally—AUC approximately 0.71—with no reported clinical utility, calibration, or treatment-interaction analysis in the abstract. ‘Prospective cohort’ does not necessarily mean prospective clinical use. The work requires complete archiving of prompts, generated code, dependencies, corrections, and model versions; otherwise reproducibility may be worse, not better. Expert supervision remained essential.
Verdict: Worth scanning for AI-methodology and education researchers; the clinical predictor itself is not deployment-ready.
6. LLM-guided reinforcement learning for lung planning
Wang Z, Guo H, Lei Y, et al. A Modular Multi-Agent Reinforcement Learning Framework Guided by LLMs: Improving Quality and Efficiency of Treatment Planning in Lung Radiotherapy. Int J Radiat Oncol Biol Phys. Published September 17, 2026.
Title: A Modular Multi-Agent Reinforcement Learning Framework Guided by LLMs: Improving Quality and Efficiency of Treatment Planning in Lung Radiotherapy
Authors: Zipai Wang, Hao Guo, Yang Lei, Robert Samstein, Kenneth E. Rosenzweig, Ming Chao, Tian Liu, Junyi Xia, Jiahan Zhang
Read the Full Article on International Journal of Radiation Oncology*Biology*Physics

Study type: Retrospective technical extension of an agentic planning platform.
Data and methods: Task-specific reinforcement-learning agents modified dose objectives under an LLM supervisory layer. The system was evaluated in 62 locally advanced NSCLC cases against institutional knowledge-based planning and an LLM-only workflow.
Principal findings
- All LLM-guided configurations achieved 98% clinical-goal compliance, versus 74% for KBP alone.
- Among 16 cases requiring refinement, the hybrid system used 75% fewer LLM calls: 1.7±0.7 versus 6.7±5.7
- Runtime decreased from 33.4±32.1 to 20.6±15.7 minutes.
- Lung-sparing mode reduced: V20: 25.2%→21.3%; p<0.001 Mean lung dose: 15.0→14.2 Gy; p<0.001
- Lung sparing caused small increases in esophageal and cord maximum doses, although constraints remained satisfied.
What changed: Modular reinforcement learning reduces reliance on repeated LLM reasoning and may make agentic planning more efficient and deterministic.
Why it matters: Separating supervision from narrow, constrained optimization actions is likely safer and more scalable than asking a general-purpose LLM to control every planning step.
Limitations and risk of bias: This uses the same 62-patient corpus as the PlanningCopilot publications covered previously and is not independent validation. Goals were institution-specific constraints rather than blinded clinical acceptability or patient outcomes. The extra lung sparing was obtained through trade-offs elsewhere. There was no new external-center, prospective-treatment, unusual-anatomy, failure-injection, or stochastic-reproducibility evaluation.
Verdict: Worth scanning as a technical advance; monitor rather than count as independent efficacy evidence.
Three most consequential developments
- Functional avoidance may become operationally scalable. AP-FLART reduced planning from hours to minutes, but the claimed pneumonitis benefit remains modeled rather than observed.
- Selective autonomy is emerging as a credible safety architecture. The on-premise agent study shows that repeated-run stability can identify a high-accuracy subset, although abstention thresholds require prospective clinical validation.
- Validation is moving closer to real translation—but sample sizes remain fragile. External-center proton planning and prospective PDAC biomarker evaluation are stronger designs than internal cross-validation, yet both remain too small to establish utility.
Implications
- Radiotherapy practice: None of this week’s studies supports unreviewed planning. Functional-lung and proton-planning systems should be commissioned with blinded plan review, dose-delivery QA, failure-case libraries, and prospective toxicity capture.
- Research design: Planning studies need active human-time measurements, multivendor testing, prespecified noninferiority margins, catastrophic-failure endpoints, and clinical rather than modeled outcomes. Biomarker studies claiming prediction must demonstrate a biomarker-by-treatment interaction, preferably within randomized data.
- Global oncology and equity: Local agents, rapid model adaptation, and GPT-assisted analytics may distribute expertise more widely. Conversely, SPECT, proton therapy, enterprise TPS integration, repeated LLM inference, and expert oversight are resource-intensive. Equity claims should therefore include total implementation cost, hardware, maintenance, connectivity, and local workforce requirements.
Actionable research opportunities
- Prospective FLART implementation trial: Compare conventional versus automated functional-avoidance planning using delivered high-function-lung dose, grade ≥2 pneumonitis, pulmonary function, quality of life, planner time, and severe plan-error rate.
- Global planning-transfer benchmark: Test locked photon and proton models across Canada, Brazil, and additional health systems, incorporating scanner/TPS variation, local constraints, atypical anatomy, blinded review, and cost per acceptable plan.
- Selective-autonomy oncology MDT study: Add repeated-run consistency and mandatory abstention to a guideline-grounded lung-cancer agent, then evaluate coverage, severe-error rate, calibration, management-changing disagreement, and the proportion of cases safely deferred to human review.”
Other posts featuring Fabio Ynoe de Moraes on OncoDaily.