用分层检索增强学习,提升眼科手术视频语言预训练效果
OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining

- 构建分层视频-文本对数据集,支持细粒度与长程视觉表征学习
- 在11个数据集上实现手术阶段识别与多器械检测的领先性能
- 适合关注医疗视觉语言模型、手术自动化研究者
外科实践涉及复杂的视觉理解、操作技能和高级医学知识,导致外科视觉-语言预训练(VLP)面临挑战,尤其受限于标注数据稀缺。为此,我们提出OphCLIP,一种专为眼科手术流程理解设计的分层检索增强型视觉-语言预训练框架。该框架利用我们构建的OphVL数据集,包含超过37.5万组分层结构化的视频-文本对,涵盖数万种不同组合属性(如手术类型、阶段/操作/动作、器械、药物,以及病因、手术目标、术后恢复建议等)。这些分层对应关系使OphCLIP能够通过短视频片段与详细描述对齐学习细粒度表示,通过完整视频与结构化标题对齐学习长期过程理解。此外,OphCLIP设计了检索增强预训练机制,自动挖掘大规模未标注手术视频中的语义相关内容,以增强叙事视频的表征学习。在11个数据集上的阶段识别与多器械识别任务中,OphCLIP展现出强大泛化能力与卓越性能。
原文摘要 · Abstract (English)
Surgical practice involves complex visual interpretation, procedural skills, and advanced medical knowledge, making surgical vision-language pretraining (VLP) particularly challenging due to this complexity and the limited availability of annotated data. To address the gap, we propose OphCLIP, a hierarchical retrieval-augmented vision-language pretraining framework specifically designed for ophthalmic surgical workflow understanding. OphCLIP leverages the OphVL dataset we constructed, a large-scale and comprehensive collection of over 375K hierarchically structured video-text pairs with tens of thousands of different combinations of attributes (surgeries, phases/operations/actions, instruments, medications, as well as more advanced aspects like the causes of eye diseases, surgical objectives, and postoperative recovery recommendations, etc). These hierarchical video-text correspondences enable OphCLIP to learn both fine-grained and long-term visual representations by aligning short video clips with detailed narrative descriptions and full videos with structured titles, capturing intricate surgical details and high-level procedural insights, respectively. Our OphCLIP also designs a retrieval-augmented pretraining framework to leverage the underexplored large-scale silent surgical procedure videos, automatically retrieving semantically relevant content to enhance the representation learning of narrative videos. Evaluation across 11 datasets for phase recognition and multi-instrument identification shows OphCLIP's robust generalization and superior performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。