arXiv:2512.11906cs.CVcs.LG2025-12被引 1

用视觉提示让大模型读懂病理切片,生成专业报告。

MPath: Multimodal Pathology Report Generation from Whole Slide Images

  • 用视觉提示将病理切片特征注入语言模型,不训练主干
  • 在竞赛中排名第四,效果接近全量训练模型
  • 适合医疗AI研究者和临床辅助诊断系统开发者

从全切片图像(WSI)自动生成诊断病理报告是计算病理学的新兴方向。由于组织形态差异大且病理描述结构复杂,将高分辨率组织模式转化为临床连贯文本仍具挑战。我们提出MPath,一种轻量级多模态框架,通过学习的视觉前缀提示机制,将预训练生物医学语言模型(BioBART)与WSI提取的视觉嵌入相结合。不同于端到端视觉-语言预训练,MPath采用基础模型的WSI特征(CONCH + Titan),通过紧凑投影模块注入BioBART,保持语言主干冻结以确保稳定性和数据效率。MPath在RED 2025 Grand Challenge数据集上开发并评估,在测试阶段2中排名第4,尽管提交机会有限。结果表明,基于提示的多模态条件化是一种可扩展且可解释的病理报告生成策略。

原文摘要 · Abstract (English)

Automated generation of diagnostic pathology reports directly from whole slide images (WSIs) is an emerging direction in computational pathology. Translating high-resolution tissue patterns into clinically coherent text remains difficult due to large morphological variability and the complex structure of pathology narratives. We introduce MPath, a lightweight multimodal framework that conditions a pretrained biomedical language model (BioBART) on WSI-derived visual embeddings through a learned visual-prefix prompting mechanism. Instead of end-to-end vision-language pretraining, MPath leverages foundation-model WSI features (CONCH + Titan) and injects them into BioBART via a compact projection module, keeping the language backbone frozen for stability and data efficiency. MPath was developed and evaluated on the RED 2025 Grand Challenge dataset and ranked 4th in Test Phase 2, despite limited submission opportunities. The results highlight the potential of prompt-based multimodal conditioning as a scalable and interpretable strategy for pathology report generation.

病理报告生成多模态视觉提示生物医学语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。