SCOUT让病理报告生成更贴近真实诊断,融合微观细胞与整体组织信息。
Semantic Context-aware mOdality fUsion Transformer (SCOUT): A Context-Aware Multimodal Transformer for Concept-Grounded Pathology Report Generation

- 用全局滑片信息和诊断概念逐步优化图像表示,实现上下文感知融合。
- 在多个数据集上超越现有方法,最高达0.865的BLEU-1和0.568的METEOR。
- 适合需要临床可解释性的病理报告自动化系统开发者使用。
全切片图像(WSIs)因极高分辨率、多尺度异质性及临床可靠性要求,给计算病理学带来根本挑战。尽管近期病理基础模型实现了流畅的报告生成,但常缺乏临床依据,无法准确体现病理科医生观察到的关键诊断概念与关系。这源于难以整合涵盖细粒度细胞模式、整片组织架构及高层诊断概念的异构视觉证据,同时保持可解释性与临床一致性。本文提出SCOUT:一种面向概念接地病理报告生成的上下文感知多模态变换器框架,通过全局滑片信息和显式诊断概念对图像表征进行渐进式条件化。该方法在统一学习范式中融合局部组织学模式、全片上下文与专家标注的语义描述,使视觉特征在编码过程中动态优化。结合深度感知上下文调制与生成过程中的自适应多模态融合,框架生成具有临床一致性的报告,并保留跨尺度表征的互补性。基于CONCH1.5特征,在TCGA-BRCA、MICCAI REG和HistAI上评估,SCOUT在所有数据集上均取得最优的BLEU-1至BLEU-4与METEOR得分,且在TCGA-BRCA和MICCAI REG上获得最佳ROUGE-L。在TCGA-BRCA上达到0.436/0.303/0.202/0.156的BLEU-1/2/3/4和0.204 METEOR;在REG 2025上实现0.865/0.834/0.805/0.780和0.568。结果验证了渐进式上下文条件化在构建临床接地报告中的有效性。
原文摘要 · Abstract (English)
Whole-slide images (WSIs) present a fundamental challenge for computational pathology due to their extreme resolution, multi-scale heterogeneity, and the requirement for clinically reliable interpretation. Although recent pathology foundation models have enabled fluent report generation, they often lack clinical grounding, failing to accurately represent key diagnostic concepts and relationships observed by pathologists. This limitation arises from the difficulty of integrating heterogeneous visual evidence spanning fine-grained cellular patterns, slide-level tissue architecture, and high-level diagnostic concepts, while maintaining interpretability and clinical coherence. Here we present SCOUT: Semantic Context-aware mOdality fUsion Transformer, a context-aware concept-grounded multimodal framework for pathology report generation that enables progressive conditioning of image representations by global slide information and explicit diagnostic concepts. The method integrates local histological patterns, whole-slide context, and expert-curated semantic descriptors within a unified learning paradigm, allowing visual features to be dynamically refined throughout the encoding process. By combining depth-aware contextual modulation with adaptive multimodal fusion during text generation, the framework produces clinically coherent reports while preserving complementarity across representational scales. Using CONCH1.5 features, we evaluate SCOUT against WSI-Caption, HistGen, and BiGen on TCGA-BRCA, MICCAI REG, and HistAI. SCOUT achieves the best BLEU-1 to BLEU-4 and METEOR scores on all datasets, plus the best ROUGE-L on TCGA-BRCA and MICCAI REG. On TCGA-BRCA, it reaches 0.436/0.303/0.202/0.156 BLEU-1/2/3/4 and 0.204 METEOR; on REG 2025, it achieves 0.865/0.834/0.805/0.780 and 0.568. These results support progressive contextual conditioning for grounded pathology report generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。