构建了最大最全的手术视觉语言数据集,助力智能手术系统理解复杂流程。
SurgLaVi: Large-Scale Hierarchical Dataset for Surgical Vision-Language Representation Learning
- 全自动流水线生成细粒度手术视频标注,分层标记粗中细三阶段流程
- 含近24万段视频-描述对,覆盖200+手术类型,质量高于现有数据集
- 适合研究手术理解、自动识别与基础模型训练的学者和工程师使用
视觉语言预训练(VLP)通过对齐语言与手术视频,在无需专家标注的情况下实现手术流程理解与任务迁移。然而,现有数据集在规模、流程多样性、语义质量与层级结构方面仍显不足。本文提出SurgLaVi,迄今最大最多样化的手术视觉语言数据集,包含近24万段视频-描述对,来自200余种手术,并具有粗、中、细三级层次结构。其核心为全自动流水线,系统生成手术视频的细粒度转录并分割为连贯流程单元;通过双模态过滤去除噪声样本,确保标注高质量且富含上下文信息。为提升可及性,我们发布SurgLaVi-η,一个由公开数据构建的11.3万对子集,规模超现有数据集四倍。为验证数据价值,我们引入SurgCLIP——一种双编码器的视频-文本对比框架。SurgCLIP在阶段、步骤、动作与器械识别任务上均显著优于现有方法,验证了大规模、语义丰富、分层结构数据对强泛化表示的直接促进作用,确立SurgLaVi作为手术基础模型研发的关键资源。
原文摘要 · Abstract (English)
Vision-language pre-training (VLP) offers unique advantages for surgery by aligning language with surgical videos, enabling workflow understanding and transfer across tasks without relying on expert-labeled datasets. However, progress in surgical VLP remains constrained by the limited scale, procedural diversity, semantic quality, and hierarchical structure of existing datasets. In this work, we present SurgLaVi, the largest and most diverse surgical vision-language dataset to date, comprising nearly 240k clip-caption pairs from more than 200 procedures, and featuring hierarchical levels at coarse-, mid-, and fine-level. At the core of SurgLaVi lies a fully automated pipeline that systematically generates fine-grained transcriptions of surgical videos and segments them into coherent procedural units. To ensure high-quality annotations, it applies dual-modality filtering to remove irrelevant and noisy samples. Within this framework, the resulting captions are enriched with contextual detail, producing annotations that are both semantically rich and easy to interpret. To ensure accessibility, we release SurgLaVi-$\b{eta}$, an open-source derivative of 113k clip-caption pairs constructed entirely from public data, which is over four times larger than existing surgical VLP datasets. To demonstrate the value of the SurgLaVi datasets, we introduce SurgCLIP, a CLIP-style video-text contrastive framework with dual encoders, as a representative base model. SurgCLIP achieves consistent improvements across phase, step, action, and tool recognition, surpassing prior state-of-the-art methods, often by large margins. These results validate that large-scale, semantically rich, and hierarchically structured datasets directly translate into stronger and more generalizable representations, establishing SurgLaVi as a key resource for developing surgical foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。