AI and Cancer Research ── Cancer biology ── 2026-10-05

Cancer molecular biology and AI
- Abstract
Surgical notes contain essential clinical information for postoperative care, yet free-text format and institutional variability limit their use for standardized data representation and secondary analysis. Biomedical entity linking enables mapping of heterogeneous clinical expressions to standardized ontologies such as SNOMED-CT, supporting semantic interoperability. However, existing approaches often rely on predefined mention spans through named entity recognition (NER), which is labor-intensive and may introduce errors. We analyzed 9,051 gastric cancer surgical notes from Seoul National University Hospital. We developed a framework that leverages an open-source large language model (LLM; LLaMA-3.1-8B) to identify contextually relevant text segments, termed evidence spans, which provide cues for ontology-based entity linking. These spans were explicitly marked and used to fine-tune SapBERT, a pretrained embedding-based biomedical encoder. We compared multiple input variants against conventional pipelines and LLM-based approaches, including in-context learning and re-ranking. Incorporating evidence spans improved entity linking performance across metrics, with gains of +2.7 in Recall@1 and +2.2 in mean Average Precision at 3 (mAP@3) compared to raw text inputs. Evidence-guided models outperformed other LLM-based approaches, with additional gains when using the evidence marker token as the pooled representation. Attention analysis indicated that explicit evidence span marking reinforced the model’s focus on ontology-relevant context while reducing attention to irrelevant text. Leveraging LLM-derived contextual evidence improves ontology-based representation of clinical text by enhancing biomedical entity linking. This approach provides a practical strategy for standardizing unstructured surgical notes and supports more reliable secondary use of clinical data in real-world healthcare settings. More broadly, the framework supports mapping of unstructured clinical text to standardized ontologies, contributing to semantic interoperability and enabling downstream secondary use of clinical data.
Journal IF-equivalent: 6.6 (OpenAlex 2-year mean citedness, value as of 2026-10-04, retrieved 2026-10-05; not the official Clarivate IF)Reference: Shin C, Eom D, You H, Kim K, Kim S, Yoon HJ. Improving clinical data standardization in surgical notes through biomedical entity linking with contextual evidence from large language models. BMC Med Inform Decis Mak. 2026 Oct 3 [Epub ahead of print]. doi:10.1186/s12911-026-03877-4.Checked: Abstract only - Abstract
Diagnosing Alzheimer’s Disease (AD), Multiple Sclerosis (MS), and brain tumors early enough to matter still depends heavily on brain MRI, and brain MRI still depends on radiologists who are in short supply almost everywhere. Convolutional neural networks have made real progress on this problem, but a single convolution only sees a small patch of the image, which limits how well a purely convolutional model can tie together pathology that is spread across the brain. This study addresses that limitation from two directions at once: a controlled, same-protocol comparison of established CNN backbones, and a lightweight attention mechanism grafted onto one of them. We classify the full eight-class Multi-Class Neurological Disorder (MCND) dataset (AD MildDemented, AD ModerateDemented, AD VeryMildDemented, Multiple Sclerosis, Normal, and three brain-tumor subtypes: glioma, meningioma, and pituitary), comprising 9,564 MRI images split 80:20 into 7,651 training and 1,913 test images. Four established CNNs, VGG-16, ResNet-50, DenseNet121, and EfficientNet-B0, were fine-tuned under one shared protocol to serve as controlled baselines. We then built VGG-16 + CBAM: a VGG16 backbone fitted with a Convolutional Block Attention Module and a much smaller classification head in place of VGG-16’s original fully connected layers, trained for 15 epochs under the same optimizer and augmentation settings used for the baselines. It reached 98.90% accuracy, a weighted F1-score of 0.9890, a macro F1 of 0.9893, a macro AUC-ROC of 0.9997, and an MCC of 0.9874 on the held-out test set, beating all four baselines (98.12%–98.85% accuracy) while using just 14.88 million parameters, roughly 9% of VGG-16’s 134.29 million, for a 56.8 MB model. We also report parameter counts, model size, FLOPs, GPU inference latency, and training time for every baseline, so the accuracy gain can be weighed against actual computational cost rather than accuracy alone. Grad-CAM maps for all eight classes show the model focusing on plausible anatomical regions: periventricular and hippocampal areas for the AD stages, white matter for MS, and the tumor core for each tumor subtype. A t-SNE projection of the learned features shows the eight classes forming distinct, mostly non-overlapping clusters. Taken together, the results suggest a small, attention-augmented CNN can match or beat larger backbones on this task at a fraction of the cost, without losing the interpretability a clinician would need to trust it.
Journal IF-equivalent: 5.2 (OpenAlex 2-year mean citedness, value as of 2026-10-04, retrieved 2026-10-05; not the official Clarivate IF)Reference: Alahmadi A, Asim N, Khan MZ, Sirshar M, Aljubayri I. An explainable vision transformer approach for automated neurological disorder classification from brain MRI scans. Sci Rep. 2026 Oct 2 [Epub ahead of print]. doi:10.1038/s41598-026-71565-4.Checked: Abstract only