AI and Cancer Research ── Lung ── 2026-10-03
Lung cancer: treatment, trials and AI
- Abstract
Background: Postoperative pulmonary complications (PPCs) significantly impair recovery after lung cancer surgery. We developed and validated an ensemble machine learning (ML) model to identify high-risk patients across the perioperative period.
Methods: We retrospectively analyzed data from 3,476 patients undergoing anatomic lung resection between 2016 and 2020 (Cohort 1). Predictors were selected through clinical expertise, statistical significance, and the Boruta algorithm. Seven ML models were developed, with the soft Voting Classifier (integrating Random Forest, CatBoost, and XGBoost) identified as the optimal approach. A prospective temporal validation cohort (n = 171) was also evaluated. Performance was assessed by AUROC with patient-level bootstrap 95% CIs, Brier score, calibration, and decision-curve analysis..
Results: The overall incidence of PPCs was 15.9%. In internal validation, the Voting Classifier ranked highest among eight models for overall PPCs (AUROC 0.7216; 95% CI 0.7044–0.7369), comparable to XGBoost and CatBoost with overlapping confidence intervals, and numerically higher than logistic regression (0.7009). In temporal validation, the model maintained moderate discriminative ability (AUROC 0.6366; 95% CI 0.5179–0.7551). Despite a surgical paradigm shift (VATS from 73.8% to 98.2%), the model retained discriminative ability for postoperative pneumonia (AUROC 0.7434; 95% CI 0.5887–0.8759), the strongest of the three subtype predictions (internal AUROC 0.7799); discrimination was inconclusive for prolonged air leak (0.6790; 0.4673–0.8796) .
Conclusions: The Voting Classifier provides an interpretable tool for predicting PPCs and guiding perioperative risk stratification, enabling targeted interventions for high-risk patients. The model was developed in patients undergoing upfront anatomic lung resection and does not apply to those receiving neoadjuvant therapy Ethical Registration ChiCTR2500102082 on May 8 2025
Journal IF-equivalent: 1.7 (OpenAlex 2-year mean citedness, value as of 2026-10-02, retrieved 2026-10-03; not the official Clarivate IF)Reference: Zhou S, Tian L, Wang X, Ma C, Li S, Cha L, et al. Ensemble Machine Learning for Predicting Postoperative Pulmonary Complications (PPCs) in Lung Surgery: A Prospective Temporal Validation and Clinical Evaluation Study. Current Problems in Surgery. 2026 Oct [Epub ahead of print]. doi:10.1016/j.cpsurg.2026.102191.Checked: Abstract only - Abstract
Background: Anemia represents a frequent and clinically burdensome complication in patients with primary lung cancer receiving platinum-based regimens, yet robust risk stratification tools that integrate multidimensional clinical data and undergo external validation remain scarce. This multicenter study sought to develop and externally validate a machine learning-based predictive model for CIA in this population.
Methods: A total of 3,434 patients with primary lung cancer were retrospectively enrolled from two geographically distinct institutions, with 3,097 assigned to the training/internal validation set (7:3 split) and 337 retained as an independent external validation set. Variables with missing rates exceeding 30% were excluded, and remaining missing data were imputed using multiple imputation by chained equations (MICE). Thirteen candidate predictors were preselected via univariate screening and least absolute shrinkage and selection operator (LASSO) regression. Nine machine learning algorithms were developed and comparatively evaluated using area under the receiver operating characteristic curve (AUC), calibration plots, and decision curve analysis. Model interpretability was achieved through SHapley Additive exPlanations (SHAP). The primary outcome was defined as hemoglobin levels < 100 g/L (CTCAE Grade ≥ 2) recorded during the first two cycles of chemotherapy.
Results: Among the nine candidate models, LightGBM achieved the highest discriminative performance, with AUCs of 0.871 (95% CI: 0.852–0.891) in the training set, 0.797 (95% CI: 0.763–0.831) in internal validation, and 0.684 (95% CI: 0.626–0.742) in the external validation set derived from Linfen Central Hospital. Calibration and decision curve analyses consistently demonstrated superior agreement and net clinical benefit for LightGBM across all three datasets. SHAP-based interpretation identified lymphocyte count, D-dimer, alanine aminotransferase, total bilirubin, sex, and albumin as the six most influential predictors, with D-dimer, female sex, and outpatient admission positively associated with anemia risk, whereas lymphocyte count, liver function indices, and albumin showed inverse associations.
Conclusions: The LightGBM-based model developed herein demonstrates acceptable and moderate predictive performance, satisfactory calibration, and potential clinical utility across geographically diverse populations, offering a preliminary risk-stratification tool for early identification of lung cancer patients at elevated anemia risk prior to chemotherapy initiation. However, given the modest external validation performance, cautious interpretation and further refinement are warranted before widespread clinical deployment.
Journal IF-equivalent: 2.7 (OpenAlex 2-year mean citedness, value as of 2026-10-02, retrieved 2026-10-03; not the official Clarivate IF)Reference: Zhang K, Li X, Zhang X, Wang X, Zhang H, Liu S, et al. Development and validation of a machine learning-based prediction model for anemia risk in patients with primary lung cancer receiving platinum-based chemotherapy: a multicenter study. Front Oncol. 2026 Sep 30;16:1948751. doi:10.3389/fonc.2026.1948751.Checked: Abstract only