Multi-Modal AI May Refine Adjuvant Escalation in High-Risk HR+/HER2− Early Breast Cancer

Multi-Modal AI May Refine Adjuvant Escalation in High-Risk HR+/HER2− Early Breast Cancer

Clinical risk remains the foundation for selecting adjuvant treatment in hormone receptor–positive, HER2-negative early breast cancer. Yet node-positive disease is biologically heterogeneous: some patients classified as high risk by conventional clinicopathologic criteria remain disease free with standard chemoendocrine therapy, while others experience recurrence despite intensive treatment.

A new prospective-retrospective analysis of the phase III UNIRAD trial explores whether artificial intelligence applied to routine pathology could help resolve this problem. The study evaluated Ataraxis Breast CTX, a causal multi-modal AI model that integrates clinical variables with features extracted from hematoxylin-and-eosin–stained tumor slides to estimate both residual recurrence risk and predicted chemotherapy benefit. The investigators asked whether the model could identify patients with excellent outcomes after standard chemoendocrine therapy and, separately, whether it could identify a subgroup more likely to benefit from additional treatment escalation.

Among 556 clinically high-risk patients with node-positive HR+/HER2− early breast cancer, the model separated a low-risk population with a 5-year disease-free survival of 93% from a high-risk population with a 5-year DFS of 80%. In an exploratory analysis, the model’s chemotherapy-benefit score also showed a statistically significant interaction with everolimus treatment, raising the possibility that AI-derived phenotypes could eventually help distinguish patients who need additional adjuvant therapy from those for whom escalation may provide little absolute benefit.

The study is provocative, but its implications require careful interpretation. CTX has not yet been prospectively validated to guide contemporary treatment decisions with ribociclib or abemaciclib, and the everolimus analysis was exploratory within a trial that was negative overall. The most important contribution of the study is therefore not a new treatment recommendation. It is a potential framework for moving adjuvant escalation beyond clinical eligibility alone toward individualized residual-risk assessment.

The Problem With Treating Every “High-Risk” Patient the Same

Patients with HR+/HER2− early breast cancer may remain at risk of recurrence for many years after diagnosis. Nodal involvement substantially increases that risk, and contemporary treatment strategies increasingly incorporate escalation beyond endocrine therapy and chemotherapy. The challenge is that the current definition of high risk is imperfect.

The manuscript notes that approximately 20% of patients with pN1 disease and 34% with pN2 disease may ultimately recur despite chemoendocrine therapy. This residual risk has driven development of additional adjuvant strategies, including CDK4/6 and mTOR inhibition. Yet major escalation trials treat broad clinicopathologic populations even though outcomes within those populations vary considerably.

This has practical consequences. Treatment intensification can reduce recurrence risk, but it also exposes large numbers of patients to additional toxicity, monitoring, treatment duration, and financial burden. If a patient who meets a conventional high-risk definition already has a very favorable prognosis after standard therapy, the absolute benefit from adding another systemic agent may be small.

Conversely, some patients may retain substantial residual risk despite receiving chemotherapy and endocrine therapy and could be the ones most likely to justify additional treatment. The clinically relevant question is therefore no longer simply:

  • Is this patient high risk?

It is increasingly:

  • How much risk remains after the treatment this patient has already received?

Multi-Modal AI

CTX Combines Digital Pathology With Clinical Information

Ataraxis Breast CTX was developed to address that distinction. The model integrates clinical variables with quantitative features derived from routine H&E whole-slide images. The pathology images are divided into smaller patches, encoded using a pathology foundation model, and combined with clinical data to produce patient-level representations.

The system generates two different outputs. CTX-prognostic estimates the patient’s predicted recurrence risk after adjuvant chemoendocrine therapy. CTX-benefit estimates predicted chemotherapy benefit by calculating the difference between modeled recurrence risk with endocrine therapy alone and with chemoendocrine therapy. This distinction is important. A prognostic model asks what is likely to happen to the patient? A treatment-effect model asks how much does a specific treatment alter that outcome? Those are fundamentally different clinical questions, and the authors attempt to model both.

The Model Was Trained Independently of UNIRAD

One methodological strength is that UNIRAD was not used to develop the model. CTX had previously been trained on 9,141 patients from multiple independent institutions. It was locked at the end of development, with no retraining or recalibration using the UNIRAD dataset.

The geographic diversity of the training population is also notable. Figure 1 shows that the development cohort included patients from North America, Europe, Asia-Pacific, and Latin America. For the UNIRAD validation analysis, investigators selected patients who had available H&E slides, had not received neoadjuvant chemotherapy, had complete clinical information, and had received adjuvant chemotherapy.

The final population included 556 patients from the original 1,278-patient UNIRAD trial. This was a clinically high-risk population: 240 patients had four or more positive nodes, 408 had tumors measuring at least 20 mm, and 180 had grade 3 disease.

AI Identified a Clinically High-Risk Group With Surprisingly Favorable Outcomes

Of the 556 patients 288 patients, or 52%, were classified as low risk and 268 patients, or 48%, were classified as high risk. After a median follow-up of approximately five years, outcomes separated substantially. Five-year DFS was:

  • 93% in the CTX low-risk group

versus

  • 80% in the CTX high-risk group.

The difference was statistically significant, with P<0.001. This is perhaps the study’s most clinically relevant result. Every patient in the analysis had already met criteria for a high-risk adjuvant clinical trial. Yet approximately half were reclassified by the AI model into a group with a 93% probability of remaining disease free at five years. The result highlights how different clinical risk at diagnosis can be from residual risk after standard therapy.

The Signal Persisted Beyond Conventional Clinicopathologic Variables

CTX-prognostic had a C-index of 0.67 for DFS in the overall cohort. A C-index of this magnitude does not represent perfect individual prediction. What is more relevant is whether the model added information beyond variables clinicians already use. In a multivariable analysis adjusting for age, tumor size, nodal involvement, grade, menopausal status, and treatment arm, CTX-prognostic remained independently associated with DFS:

HR 1.31 per standard-deviation increase; 95% CI, 1.06–1.61; P=0.012. An exploratory analysis incorporating additional pathological factors—including lymphovascular invasion, ER and PR expression, HER2 staining, and ductal or lobular histology, again found CTX-prognostic independently associated with DFS.

The model therefore appears to capture biological information that is not completely represented by standard staging and pathology variables.

Risk Stratification Extended to Distant Recurrence and Overall Survival

The prognostic signal was not limited to the primary DFS endpoint. At five years, distant metastasis-free survival was:

  • 93% in low-risk patients

versus

  • 82% in high-risk patients.

Five-year overall survival was:

  • 98% versus 92%.

When modeled continuously and adjusted for clinical covariates, CTX-prognostic remained associated with both distant metastasis-free survival and overall survival. These findings strengthen the argument that the model is capturing clinically meaningful residual disease biology rather than merely predicting local or less consequential events.

The Most Provocative Finding Came From Everolimus

UNIRAD originally evaluated whether adding everolimus to adjuvant endocrine therapy could improve outcomes in high-risk HR+/HER2− early breast cancer. The overall trial was negative. That history makes the new exploratory analysis particularly interesting. The investigators asked whether CTX-benefit—a score originally developed to estimate chemotherapy responsiveness, could identify a subgroup in whom the effect of everolimus differed.

Among patients classified as having low predicted chemotherapy benefit, five-year DFS was:

  • 95% with everolimus

versus

  • 85% in the control arm.

The difference was statistically significant by log-rank testing. By contrast, among patients classified as having high predicted chemotherapy benefit, there was no statistically significant evidence that everolimus improved DFS. More importantly, formal interaction testing showed that the relationship between everolimus treatment and DFS differed according to the CTX-benefit score. In the multivariable model, the treatment-by-CTX-benefit interaction was:

  • HR 1.68; 95% CI, 1.12–2.53; P=0.01.

Multi-Modal AI

Why Would a Chemotherapy-Benefit Score Predict Everolimus Benefit?

The authors propose an interesting biological interpretation. Patients predicted to respond strongly to chemotherapy may have disease whose residual risk is substantially reduced by cytotoxic treatment. Patients predicted to derive little chemotherapy benefit, however, may retain greater residual biological risk after chemotherapy and therefore be more likely to benefit from treatment targeting an alternative pathway.

In this framework, CTX is not simply looking for patients with the highest baseline risk. It attempts to identify what kind of risk remains after a given therapy. The authors suggest that the same conceptual approach could ultimately be relevant to contemporary escalation with agents such as CDK4/6 inhibitors. That idea is compelling but remains hypothetical.

The study did not evaluate ribociclib or abemaciclib, and the everolimus finding cannot be assumed to predict benefit from CDK4/6 inhibition.

The Histology Maps Suggest the AI Is Seeing Biologically Plausible Features

One concern with complex AI models is whether their predictions reflect biologically meaningful information or opaque statistical correlations. The investigators therefore generated spatial score maps and had breast pathologists review representative cases while blinded to clinical outcomes.

The model’s phenotypes corresponded with recognizable morphology. Low-risk/low-benefit tumors were often well differentiated, mucinous, and characterized by prominent in situ disease with relatively little invasive carcinoma. High-risk/high-benefit tumors tended to show high cellularity, prominent mitotic activity, infiltrative tumor cords, lymphocytic inflammation, fat involvement, and poorly differentiated morphology.

High-risk/low-benefit tumors showed a different phenotype, including intermediate cellularity, mucin, and prominent in situ components.

Could AI Identify Patients Who Do Not Need CDK4/6 Escalation?

This is where the study becomes most relevant to current practice. The authors note that all patients in the UNIRAD cohort would meet criteria for adjuvant ribociclib, yet the CTX low-risk population had a 5-year DFS of 93% despite having been selected clinically as high risk.

For context, the manuscript cites an approximately 4.5% absolute 5-year invasive DFS improvement with ribociclib in NATALEE and a 7.6% absolute improvement with abemaciclib in monarchE. That raises a legitimate clinical question.

If a subgroup already has a very favorable residual prognosis after chemotherapy and endocrine therapy, could the absolute benefit of another two or three years of systemic treatment be sufficiently small that escalation becomes unnecessary? The current study cannot answer that. CTX was not tested as a treatment-selection biomarker in NATALEE or monarchE, and no randomized comparison shows that CTX-low patients can safely omit CDK4/6 inhibition.

Therefore, the 93% DFS figure should be considered a rationale for prospective de-escalation research, not evidence to withhold established therapy from an otherwise eligible patient today.

Prognostic and Predictive Biomarkers Must Not Be Confused

This distinction is essential. CTX-prognostic clearly demonstrated the ability to separate groups with different outcomes. That makes it a prognostic biomarker in this dataset. The everolimus analysis provides evidence of a treatment interaction with CTX-benefit and therefore raises a predictive hypothesis. However, prediction of treatment effect requires a higher evidentiary standard than prognosis.

The original UNIRAD trial did not show a significant overall DFS benefit from everolimus, and the subgroup analysis in this manuscript was exploratory. It therefore needs independent validation before it could be used to select patients for everolimus or extrapolated to other escalation therapies.

A Negative Trial May Contain a Positive Subgroup, but That Requires Validation

The everolimus finding illustrates a broader problem in oncology drug development. A therapy can fail to improve outcomes across an entire clinically defined population if only a biologically distinct subset is treatment sensitive. The authors suggest that this may partly explain why both UNIRAD and SWOG 1207 failed to establish adjuvant everolimus as a standard strategy despite the drug’s established activity in metastatic HR+/HER2− breast cancer.

But retrospective biomarker discovery within a negative trial is vulnerable to false-positive findings. For that reason, the appropriate next step is not clinical adoption but external validation, ideally in another randomized dataset such as SWOG 1207. The manuscript explicitly identifies this as an important future direction.

Limitations Are Important

Several limitations substantially affect interpretation. First, only 56% of the original UNIRAD population had H&E slides available, and modest differences existed between patients with and without available tissue in T stage, nodal status, grade, and menopausal status.

Second, patients receiving neoadjuvant chemotherapy were excluded because treatment could alter tumor morphology and thereby affect the image-based model. CTX performance therefore remains unknown in the increasingly common post-neoadjuvant setting. Third, everolimus had a high treatment-discontinuation rate in UNIRAD, which may influence estimates of treatment effect.

Fourth, contemporary escalation strategies have changed substantially since UNIRAD. The clinically relevant question today concerns drugs such as abemaciclib and ribociclib, not routine adjuvant everolimus. Finally, the model was developed and supported by Ataraxis AI. Several authors are equity holders in the company. These relationships are appropriately disclosed and should be considered when interpreting technology-validation research.

The Study Points Toward a Different Kind of Precision Oncology

Breast cancer precision medicine has traditionally relied heavily on genomic assays. CTX represents a different approach. Instead of measuring a predefined gene-expression signature, it extracts quantitative information from conventional pathology images and combines those features with clinical variables to estimate counterfactual outcomes under different treatment strategies.

This could have practical advantages if validated. H&E slides are already generated for essentially every invasive breast cancer. A reliable model built on existing pathology could potentially deliver additional prognostic and predictive information without requiring another tissue-consuming molecular assay.

But the key word is validated. A model may discriminate risk well and still fail to improve patient outcomes when used to make treatment decisions. The ultimate test will therefore require prospective clinical utility studies showing that AI-guided escalation or de-escalation is safe and clinically beneficial.

Multi-Modal AI

The Bottom Line

This UNIRAD analysis suggests that multi-modal AI may help refine one of the most difficult decisions in HR+/HER2− early breast cancer: which clinically high-risk patients actually retain enough residual risk to justify additional adjuvant treatment?

Among 556 node-positive patients treated with adjuvant chemotherapy and endocrine therapy, CTX classified approximately half as low risk and half as high risk.

At five years:

  • DFS: 93% vs 80%
  • DMFS: 93% vs 82%
  • OS: 98% vs 92%

for low-risk versus high-risk patients.

The model remained prognostic after adjustment for standard clinicopathologic factors. Even more provocatively, an exploratory treatment-interaction analysis suggested that patients predicted to derive little benefit from chemotherapy experienced greater benefit from adjuvant everolimus, with a significant CTX-benefit–treatment interaction.

These findings do not justify withholding current standard adjuvant CDK4/6 inhibitors or selecting everolimus based on CTX today. What they do demonstrate is a potentially important principle, clinical high risk and residual biological risk are not the same thing.

The next generation of adjuvant precision medicine may therefore need to move beyond determining who is eligible for treatment escalation and toward identifying who is likely to benefit enough from escalation to justify it.

Reference

  1. Bachelot T, Chabaud S, Lemonnier J, Cottu PH, Dalenc F, Howard FM, Pusztai L, Tang C, Biswas D, Zeng K, Witowski J, Geras KJ, André F, Penault-Llorca FM. A Causal Multi-modal AI Model Stratifies Residual Risk and Identifies Candidates for Treatment Escalation in Node-Positive HR+/HER2− Early Breast Cancer.
Sona Karamyan
Fact checked by Sona Karamyan MD, Medical Oncologist
Amalya Sargsyan
Medically reviewed by Amalya Sargsyan MD, Medical Oncologist