用语言描述对齐视觉,提升胶囊内镜异常检测的泛化能力
CapCLIP: A Vision-Language Representation Alignment Approach for Wireless Capsule Endoscopy Analysis
- 通过标准术语文本对齐内镜图像,学习语义可解释的特征表示
- 零样本测试中跨数据集表现优于现有模型,尤其在图文匹配任务上提升显著
- 适合需要跨中心迁移、少标注场景的医学影像分析研究者
无线胶囊内镜(WCE)可非侵入式评估小肠,但每例检查产生大量图像帧,且在成像条件多变时难以识别细微病变。现有基于学习的方法多为纯视觉模型,通常仅针对有限病种,跨数据集和中心的迁移能力弱。本文提出专用于WCE的视觉-语言表征对齐框架CapCLIP,将内镜图像与基于标准化命名法和病灶感知模板生成的临床文本描述对齐,从而学习兼具语义信息与可迁移性的嵌入表示。在未见数据集上,以严格零样本条件评估其在三个下游任务中的表现:K近邻分类、CLIP式图文分类和文本到图像检索。结果表明,CapCLIP在各项任务中持续优于对比基线,尤其在跨分布数据集的零样本图文分类和跨模态检索中表现突出。研究证明语言引导的表示学习能增强WCE分析的泛化性和语义可解释性,为面向胶囊内镜的通用模型奠定基础。
原文摘要 · Abstract (English)
Wireless capsule endoscopy (WCE) enables non-invasive visual assessment of the small bowel, but its clinical utility is constrained by the large volume of frames generated per examination and the difficulty of recognising subtle abnormalities under highly variable imaging conditions. Existing learning-based approaches for WCE are predominantly vision-only, often confined to narrow pathology sets, and show limited transfer across datasets and centres. To address these limitations, this study introduces CapCLIP, a domain-specific vision-language representation learning framework for WCE. CapCLIP aligns capsule endoscopy frames with clinically grounded textual descriptions derived from standardised nomenclature and pathology-aware caption templates, thereby learning embeddings that are both semantically informed and transferable. The proposed framework is evaluated against relevant open-source vision and vision-language foundation models under strict zero-shot conditions using unseen WCE datasets. Evaluation covers three downstream tasks: K-nearest neighbour classification, CLIP-style image-text classification, and text-to-image retrieval. Across these settings, CapCLIP consistently outperforms the compared baselines, with particularly strong gains in zero-shot image-text classification and cross-modal retrieval on out-of-distribution datasets. The results indicate that language-guided representation learning can improve both generalisation and semantic interpretability in WCE analysis. These findings position CapCLIP as a step toward foundation models tailored to capsule endoscopy and support the use of language-grounded WCE analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。