通过三阶段自监督学习,提升结肠镜图像与文本的对齐精度。
Endo-CLIP: Progressive Self-Supervised Pre-training on Raw Colonoscopy Records
- 分三步净化数据:去背景、提取临床特征、用患者级注意力解决多病灶歧义。
- 在零样本和少样本下,结肠息肉检测与分类性能超越现有方法。
- 适合医疗视觉预训练、内镜分析研究者使用。
在图像-文本结肠镜记录上进行预训练,有望显著提升内镜图像分析能力,但面临非信息性背景图像、复杂医学术语及多病灶描述模糊等挑战。本文提出端到端自监督框架Endo-CLIP,包含三个阶段:清除、调谐与统一。首先剔除背景帧;其次利用大语言模型提取临床属性,支持细粒度对比学习;最后采用患者级交叉注意力机制解决多息肉描述歧义问题。大量实验表明,Endo-CLIP在零样本和少样本条件下,于结肠息肉检测与分类任务上均显著优于当前最优预训练方法,为更精准、更符合临床需求的内镜分析铺平道路。
原文摘要 · Abstract (English)
Pre-training on image-text colonoscopy records offers substantial potential for improving endoscopic image analysis, but faces challenges including non-informative background images, complex medical terminology, and ambiguous multi-lesion descriptions. We introduce Endo-CLIP, a novel self-supervised framework that enhances Contrastive Language-Image Pre-training (CLIP) for this domain. Endo-CLIP's three-stage framework--cleansing, attunement, and unification--addresses these challenges by (1) removing background frames, (2) leveraging large language models to extract clinical attributes for fine-grained contrastive learning, and (3) employing patient-level cross-attention to resolve multi-polyp ambiguities. Extensive experiments demonstrate that Endo-CLIP significantly outperforms state-of-the-art pre-training methods in zero-shot and few-shot polyp detection and classification, paving the way for more accurate and clinically relevant endoscopic analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。