从28万份报告中训练出可精准匹配肠镜图像与描述的模型
A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports

- 从28万份报告中提取12.6万组病变图像-文本对进行训练
- 在良恶性分类任务中接近12位专家水平,零样本表现优异
- 让医生用语言描述即可实现精准医学目标,无需逐帧标注
尽管常规肠镜报告包含丰富的专家描述,但视觉语言模型在肠镜中仍应用不足。这些报告记录病变外观、大小和位置,但整体总结整个操作过程,而非为单帧图像生成描述,导致临床发现与对应图像关联较弱。本文开发了EndoCLIP,一个基于125,756组病变级图像-文本对训练的肠镜视觉语言基础模型,该数据从280,476份常规肠镜记录中逐步恢复。在病变级图像-文本检索、结构化报告生成及六个多中心临床分类任务中,EndoCLIP在零样本和线性探测设置下均优于通用和生物医学视觉语言编码器。在良恶性分类任务中,其线性探测性能接近盲法评估中12名内镜医师的水平。结果表明,恢复病灶与图像的对应关系,可将常规文档转化为可扩展的监督信号,使临床目标可通过语言指定,而无需为每项任务单独标注。
原文摘要 · Abstract (English)
Vision-language models remain underused in colonoscopy despite the rich expert descriptions recorded in routine reports. These reports document lesion appearance, size and location but summarise entire procedures rather than caption individual frames, leaving clinical findings only weakly linked to the corresponding images. Here we develop EndoCLIP, a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs progressively recovered from 280,476 routine colonoscopy records. Across lesion-level image-text retrieval, structured report generation and six multi-centre clinical classification tasks, EndoCLIP outperforms general-purpose and biomedical vision-language encoders in both zero-shot and linear-probe settings. On benign-versus-malignant classification, its linear probe approaches the performance of expert readers in a blinded study involving 12 endoscopists. These results suggest that recovering finding-to-frame correspondence can transform routine documentation into scalable supervision, enabling clinical targets to be specified in language rather than separately annotated for each task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。