用临床报告指导视觉聚焦,让内镜模型更懂解剖结构。
EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment

- 用报告文本作为查询,稀疏聚合关键图像帧
- 在348K例数据上实现超越现有模型的性能
- 适合需要零样本泛化的医学影像分析场景
基础模型的发展对提升内镜图像分析至关重要。然而,现有内镜基础模型主要依赖单模态图像或视频的自监督学习,忽略了临床报告中的丰富语义信息。同时,由于结构化解剖描述与冗余、未标注的视觉流之间存在根本性模态鸿沟,有效利用这些记录仍具挑战。本文提出EndoVLM,一个在超过348,000例内镜检查数据上预训练的视觉-语言基础模型,每例包含一份临床报告与对应的图像集合。其采用解剖引导的稀疏池化机制,以文本描述为查询驱动稀疏注意力,高效将语义显著帧聚合为特定解剖区域的视觉表征。随后,通过渐进式语义感知对齐策略,利用结构化软标签建模解剖与病理状态,实现从全局患者级匹配到细粒度定位对齐。最后,仅对这些语义丰富的帧应用语义聚焦的掩码自编码器,融合低层视觉精度与高层语义鲁棒性。大量下游任务实验证明,EndoVLM优于现有基础模型,并保持与专用方法相当的竞争力。尤为突出的是,其展现出强大的零样本泛化能力,凸显其在更广泛临床场景中的潜力。
原文摘要 · Abstract (English)
The development of foundation models (FMs) is crucial for advancing endoscopic image analysis. However, existing endoscopy FMs mainly rely on self-supervised learning from uni-modal images or videos, overlooking the rich semantic knowledge contained in clinical reports. Furthermore, effectively leveraging these records is hindered by a fundamental modality gap: structured anatomical descriptions are not naturally mapped to specific frames within the high-redundancy, uncurated visual streams. In this paper, we present EndoVLM, a novel vision-language FM pre-trained on over 348K endoscopic examinations, each pairing a clinical report with its corresponding image collection. An Anatomy-Guided Sparse Pooling mechanism utilizes textual descriptions as queries to drive sparse attention, efficiently aggregating semantically salient frames into anatomy-specific visual representations across redundant image-sets. Next, a Progressive Semantic-Aware Alignment strategy models clinical taxonomy (anatomy and pathological status) via structured soft targets, bridging the gap from global patient-level matching to fine-grained localized alignment. Finally, a Semantic-Concentrated Masked Autoencoder is applied exclusively to these semantic-rich frames, integrating low-level visual precision with robust high-level semantic representation. Extensive experiments across various downstream tasks demonstrate that EndoVLM outperforms existing foundation models and remains competitive with task-specific methods. Remarkably, EndoVLM also exhibits robust zero-shot generalization capabilities, highlighting its potential for broader clinical application.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。