AI and Cancer Research ── Liver ── 2026-10-10

Liver cancer: treatment, trials and AI
- Abstract
Background/Objectives: Treatment allocation for hepatocellular carcinoma (HCC) is a highly complex task that requires multidisciplinary tumor board (MDT) input; however, its accessibility and reliability can vary across the healthcare system. Large language models (LLMs) have emerged as potential clinical decision-making tools to aid MDTs. Our aim was to compare different LLM-generated treatment recommendations for HCC cases with MDT decisions.
Methods: We retrospectively analyzed 100 HCC cases discussed during MDT meetings in a tertiary-care hospital in Cluj-Napoca, Romania. Identical prompts and structured clinical information were offered to four different LLMs (ChatGPT-5, a customized Tumor Board ChatGPT, Gemini 2.5 Flash, and Gemini 2.5 Pro), each of which was required to provide treatment recommendations. Concordance with the MDT decisions was assessed across the first, second, and third treatment options. Inter-model differences were evaluated using Cochran’s Q and Holm-adjusted exact McNemar tests. Chance-corrected agreement was assessed using Cohen’s kappa (κ), and generalized estimating equations evaluated associations between concordance and clinical complexity.
Results: First-recommendation concordance was 81% (95% CI 72.2–87.5) for GPT-5, 80% (71.1–86.7) for Gemini 2.5 Pro, 78% (68.9–85.0) for TumorBoard ChatGPT, and 63% (53.2–71.8) for Gemini 2.5 Flash; cumulative top-3 concordance reached 95%, 92%, 95%, and 86%, respectively. GPT-5 (κ = 0.754) and Gemini 2.5 Pro (κ = 0.743) showed substantial chance-corrected agreement. Gemini 2.5 Flash showed significantly lower concordance, including after adjustment for BCLC stage, Child–Pugh class, tumor burden, and treatment category (p = 0.006), while these clinical complexity variables were not significantly associated with concordance.
Conclusions: LLMs showed moderate-to-substantial agreement with MDT decisions, with meaningful inter-model differences, for a complex disease, such as HCC. Our study highlights that, while the need for human oversight is essential, LLMs could serve as supportive tools for expert guidance in some clinical settings and underscore AI’s potential in enhancing liver cancer care.
Journal IF-equivalent: 2.1 (OpenAlex 2-year mean citedness, value as of 2026-10-09, retrieved 2026-10-10; not the official Clarivate IF)Reference: Grapa C, Mocan T, Leucuta DC, Mocan LP, Craciun R, Stefanescu H, et al. Concordance Between Large Language Models and Multidisciplinary Tumor Board Recommendations in Treatment Allocation for Hepatocellular Carcinoma. Livers. 2026 Oct 8;6(5):104. doi:10.3390/livers6050104.Checked: Abstract only - Abstract
INTRODUCTION: This study assessed whether Gd-DTPA-enhanced MRI could predict the immunoscore in HCC noninvasively before therapy, without the need for tissue sampling. METHODS: Retrospectively, surgically treated HCC patients who had a preoperative Gd-DTPA-enhanced MRI exam were enrolled. Eligible patients were randomly allocated to a training set and a validation set following a 3:1 randomization scheme. Immunohistochemistry was performed to quantify CD3-positive and CD8-positive cell densities; patients were stratified into high and low immunoscore groups using the median cutoff. Volumes of interest encompassing hepatic lesions, including intratumoral and peritumoral 10-mm margins, were manually delineated on multiparametric MRI sequences for radiomics feature extraction. Clinical and pathological data were retrieved from electronic medical records. A Support Vector Machine (SVM) was used to develop three predictive models: (1) Intratumoral Radiomics Model (IRM), (2) combined intra- and peritumoral radiomics model (CRM), and (3) Clinical Combined Radiomics Model (CCRM). Model performance was compared using the DeLong test and evaluated via the area under the receiver operating characteristic curve (AUC), calibration curves, and Decision Curve Analysis (DCA). RESULTS: Of 111 eligible patients, 83 were assigned to the training cohort and 28 to the validation cohort. Baseline characteristics were balanced between cohorts. The CCRM demonstrated the highest AUC values in both cohorts (training: 0.947; validation: 0.913). Compared with IRM, CCRM showed significantly better performance in the validation cohort (p = 0.016) but not in the training cohort (p = 0.084). In the validation cohort, CCRM did not significantly outperform CRM (p = 0.55), whereas a significant difference was observed in the training cohort (p = 0.036). Calibration curves demonstrated good agreement. DCA indicated that the CCRM and the CRM achieved similar overall net benefits, which were higher than those of the IRM across threshold probabilities of 46% to 98%. DISCUSSION: This radiomics approach suggests the feasibility of using routine Gd-DTPA-enhanced MRI for preoperative immunoscore prediction in HCC, providing preliminary evidence for tumor immune characterization. CONCLUSION: A radiomics model integrating intra- and peritumoral MRI features enables noninvasive preoperative immunoscore prediction in HCC, which may provide preliminary support for individualized immunotherapy decision-making, although external validation is warranted.
Journal IF-equivalent: 1.3 (OpenAlex 2-year mean citedness, value as of 2026-10-09, retrieved 2026-10-10; not the official Clarivate IF)Reference: Wang J, Jiang Y, Li Z, Xu D, Teng H, Meng H, et al. Radiomics Model Based on Gd-DTPA-enhanced Multiparametric MRI Supports Noninvasive Pretreatment Prediction of Immunoscore in Hepatocellular Carcinoma. CMIR. 2026 Sep 30;22. doi:10.2174/0115734056504350260925115431.Checked: Abstract only