1. The Problem: Walls Around Intraoperative Pathology
Mid-surgery, the surgeon sends excised tissue to the pathologist. The pathologist prepares a frozen section, examines it under a microscope, and returns a judgment while the surgery waits. The entire operating team holds until the answer comes back. This intraoperative pathology is a linchpin of surgical precision, but it faces two persistent difficulties.
First, frozen sections are lower in quality than formalin-fixed specimens, making diagnosis harder. Ice crystal artifacts, tissue distortion, and staining inconsistencies all contribute to a higher error rate compared to permanent sections processed after surgery. Second, large-scale datasets of frozen sections have been scarce. Most computational pathology research has focused on formalin-fixed, paraffin-embedded specimens, which are far more abundant in digital pathology archives. The authors note that the lack of large-scale, prospective validation has impeded routine adoption of computational pathology in surgical workflows.
2. The Proposal: CRISP, a Foundation Model Built for the Operating Room
CRISP (Clinically-oriented Robust Intraoperative Support for Pathology) is a foundation model specifically designed for intraoperative pathology. It was trained on over 100,000 frozen sections collected from ten medical centers.
This paper was peer-reviewed and published in Nature Medicine. It is not a preprint. That said, passing peer review does not guarantee that every claim will replicate in independent studies.
3. What Was Shown: Within the Bounds of the Abstract
The authors report two kinds of evaluation.
Retrospectively, CRISP was evaluated on more than 15,000 intraoperative slides across nearly 100 diagnostic tasks, including benign-malignant discrimination, key intraoperative decision-making, and pan-cancer detection. The model showed stable generalization across 6 institutions, 14 tumor types, and 24 anatomical sites, including sites and rare cancers not seen during training.
Prospectively, a cohort of over 3,000 patients was used. CRISP sustained high diagnostic accuracy under real-world conditions, directly informing surgical decisions in 92.6% of cases. Human-AI collaboration reduced diagnostic workload by 35%, avoided 105 ancillary tests, and detected micrometastases with 87.5% accuracy.
4. What This Looks Like in Practice
A surgeon excises a lung margin and sends it to the pathologist to check whether cancer cells remain at the edge. A frozen section is prepared. The pathologist examines it under the microscope. If positive, the surgeon proceeds to additional resection. If negative, they begin closure.
The time this judgment takes governs the flow of surgery. With CRISP in the loop, the AI returns an initial assessment before the pathologist views the slide. The pathologist then confirms using their own eyes, with the AI's output as a reference. According to the authors, this collaborative setup reduced the pathologist's workload compared to working alone and allowed some ancillary tests to be skipped.
5. What Is Not New, and Where the Limits Lie
Foundation models for computational pathology are not new; several research groups have built them before CRISP. What distinguishes CRISP is its specialization in frozen sections and the scale of its prospective validation, which is unusual for intraoperative AI.
Limits remain. First, the abstract does not specify which regions or patient populations the ten centers represent. If the centers cluster geographically, generalization claims may be affected. Second, the 92.6% figure refers to cases where CRISP "directly informed surgical decisions," but the abstract does not define what "informed" means operationally. Whether this means the pathologist accepted the AI's judgment or the AI's judgment matched the final diagnosis are different things. Third, 87.5% accuracy in detecting micrometastases is high, but the clinical impact of the cases that were missed is a separate question.
6. How Independent Outlets Covered This Paper
Two independent outlets were identified covering this paper or its immediate context.
Nature's npj Digital Medicine ran the headline "Specialized foundation models for intelligent operating rooms"[2], framing CRISP within a broader narrative about AI-enabled operating rooms. On medRxiv, a separate model called SIGNAL was reported under "Scalable, Real-World Model for Rapid Intraoperative Molecular Classification of Gliomas Using Stimulated Raman Histology"[3], addressing an adjacent topic of intraoperative molecular classification in the same period.
The gap worth noting is one of framing. The CRISP paper focuses specifically on pathology diagnosis, while the Nature coverage headline uses the broader frame of "intelligent operating rooms." The paper's claims and the coverage headline operate at different levels of granularity. This kind of frame expansion is common in science communication: a paper about one specific application gets absorbed into a broader narrative that the paper itself does not claim. Readers encountering the coverage headline alone might expect a system that orchestrates the entire operating room, which is not what CRISP does.
7. What This Connects to in Pharma and Regulation
Intraoperative AI pathology is not directly within the business scope of pharmaceutical companies. But two connection points exist.
First, there is an intersection with companion diagnostics. If AI-driven intraoperative assessment improves the precision of surgical margins, it may influence post-surgical treatment selection. Specifically, a future where companion diagnostics for molecularly targeted therapies interface with real-time intraoperative assessment would touch pharma's biomarker strategy.
Second, this paper sets a reference point for the evidence bar in AI medical device regulation. A large-scale prospective clinical validation of an intraoperative AI model is precisely the kind of evidence that regulators look for. When regulators evaluate AI-based diagnostic tools, they look for prospective data collected under real clinical conditions, not just retrospective analysis of archived slides. CRISP's prospective cohort of over 3,000 patients, collected across multiple institutions, represents the kind of evidence package that regulatory submissions for AI medical devices will increasingly need to include. Similar AI medical devices seeking approval in the future may find this scale of prospective validation cited as a baseline expectation.