AI and Cancer Research ── Gastric ── 2026-10-10

Gastric cancer: treatment, trials and AI
- Abstract
Objective: To construct various radiomics models based on CT venous-phase images for the quantitative assessment of gastric cancer (GC) invasion depth into the gastric wall, and to compare the diagnostic performance of different machine-learning models in distinguishing early (T1–T2 stage) from advanced (T3–T4 stage) GC, aiming to identify a robust and interpretable model and evaluate its potential for clinical application.
Methods: In this retrospective study, 223 pathologically confirmed GC patients (66 early-stage, 157 advanced-stage) treated between January 2022 and May 2025 were enrolled. Patients were allocated into a training set ( n = 156) and an internal hold-out test set ( n = 67) through stratified random sampling at a 7:3 ratio. All patients underwent enhanced CT within one week prior to surgery. Venous-phase images were selected, and three-dimensional regions of interest (ROIs) encompassing the tumor were manually delineated using 3D Slicer for radiomics feature extraction. To rigorously exclude data leakage, all preprocessing steps—including feature standardization, LASSO feature selection, hyperparameter tuning, and threshold determination—were performed exclusively within the training set using a pipeline-based approach with 5-fold cross-validation. Feature selection stability was further assessed using 50 iterations of repeated stratified cross-validation. Four machine learning models were constructed: Logistic Regression (LR), Random Forest (RF), XGBoost (XGB), and Support Vector Machine (SVM). Performance was evaluated using ROC curves, PR curves, confusion matrices, and decision curve analysis (DCA). SHAP analysis was applied to the LR model for interpretability.
Results: A total of 107 radiomics features were extracted, from which 23 key features were selected by LASSO regression. Feature selection stability analysis revealed that 2 features were consistently selected across all 50 iterations (100% frequency), 1 feature was retained in ≥ 98% iterations; overall 6 features exhibited selection frequency ≥ 80%, and 8 features achieved frequency ≥ 50%, while the remaining features showed considerable selection variability. In the internal hold-out test set, the LR model achieved the AUC of 0.926 (95% CI: 0.860–0.978), followed by SVM (AUC = 0.915, 95% CI: 0.848–0.973), RF (AUC = 0.893, 95% CI: 0.814–0.960), and XGBoost (AUC = 0.889, 95% CI: 0.806–0.956). DeLong’s test revealed no statistically significant differences in AUC among the four models (all P > 0.05); the marginal difference between LR and RF (AUC difference = 0.033, P = 0.051) only represented a weak numerical gap without statistical superiority of LR. The LR model demonstrated balanced performance with accuracy 0.836, sensitivity 0.830, specificity 0.850, PPV 0.929, NPV 0.680, F1-score 0.876, and average precision (AP) 0.972; notably, LR yielded the lowest false-negative rate (17.0%), which minimizes the clinical risk of undertreating advanced GC. Hosmer-Lemeshow test revealed good calibration for the Random Forest ( P = 0.586), XGBoost ( P = 0.081), and SVM ( P = 0.213) models, whereas the Logistic Regression model showed evidence of miscalibration ( P < 0.001). Despite this calibration limitation, the LR model’s excellent discriminative performance (AUC = 0.926) supports its clinical utility as a binary classifier. DCA indicated that all four models provided clinical utility, with the LR model offering favorable net benefit across most threshold probabilities. SHAP analysis confirmed that the contribution directions of selected features such as LeastAxisLength and SurfaceVolumeRatio were consistent with pathological interpretations.
Conclusion: All four machine learning models constructed based on CT venous-phase images showed comparable discriminative power, with no statistically significant inter-model AUC differences detected via DeLong test (all P > 0.05). Rather than possessing statistically superior predictive efficacy, the LR model had higher clinical translation potential due to its excellent interpretability, balanced classification metrics and minimal false-negative risk. The LR model, with transparent coefficient structure and good diagnostic efficacy, can serve as a quantitative auxiliary tool for preoperative staging and treatment decision-making, providing an objective basis for precise GC diagnosis and treatment. Further multi-center external prospective validation is mandatory before formal clinical deployment.
Journal IF-equivalent: 5.3 (OpenAlex 2-year mean citedness, value as of 2026-10-02, retrieved 2026-10-03; not the official Clarivate IF)Reference: Zhang H, Liu Q, Lian W, Huang R, Zhao X, Li L, et al. CT venous-phase radiomics for preoperative differentiation of early and advanced gastric cancer: construction, comparison, and validation of machine-learning models. Cancer Imaging. 2026 Oct 8 [Epub ahead of print]. doi:10.1186/s40644-026-01132-7.Checked: Abstract only