构建首个大规模开放与微创手术视频-语言数据集,支持手术智能分析
SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery

- 基于公开YouTube内容构建2391小时手术视频,覆盖18个专科5000+术式
- 首次实现开放手术大规模标注,提供标准化评估基准和专家验证的问答对
- 多层级标注体系支持手术理解、推理与基础模型训练,适合临床AI研究者使用
我们提出SurgAtlas,目前规模最大、涵盖最广的外科视频-语言数据集,包含15,291段视频(总计2,391小时),覆盖18个外科专业及超过5,000种手术类型,全部来自公开可获取的YouTube内容。SurgAtlas是首个以大规模方式包含开放手术的视频-语言数据集,其中包含6,182段开放手术视频和超过9,000段微创手术记录,并首次建立针对开放手术视频理解的标准化基准。我们还提供经专家验证的子集,包含跨多种术式的视觉问答对,作为临床可信的手术推理基准。相比现有数据集,SurgAtlas具备最多样化的标注模式,融合片段级描述、步骤与阶段级说明、视频级手术描述以及基于分层分类体系的推理型问答对。这些标注通过基于LLM的自动化多级流程和具显性语义锚定验证的VQA生成框架完成。其规模与多样性使我们能够训练覆盖广泛的外科基础模型:采用两阶段指令微调流程对Qwen3-VL-8B进行微调,在多个既有基准上取得竞争性或领先性能,涵盖阶段识别、三元组检测与推理问答任务。更广泛而言,SurgAtlas提供了可用于未来大规模预训练的原生公共视频语料库,推动下一代外科多模态基础模型的发展。
原文摘要 · Abstract (English)
We introduce SurgAtlas, the largest surgical video-language dataset to date, comprising 15,291 videos (2,391 hours) spanning 18 surgical specialties and over 5,000 procedure types, sourced entirely from publicly available YouTube content. SurgAtlas is also the first surgical video-language dataset to include open surgery at scale, with 6,182 open procedure videos alongside over 9,000 minimally invasive recordings, and the first to establish standardized benchmarks for open-surgery video understanding. We additionally provide an expert-validated subset with verified visual question-answer pairs across diverse open and minimally invasive procedures, serving as a clinically grounded benchmark for surgical reasoning. Compared with existing surgical video-language datasets, SurgAtlas provides one of the most diverse annotation schemas, combining segment-level captions, step- and phase-level descriptions, video-level surgical descriptions, and reasoning-oriented question-answer pairs organized within a hierarchical taxonomy. These annotations are constructed through an automated multi-tier pipeline with LLM-based enrichment and a staged VQA generation framework with explicit groundedness verification. The scale and diversity of SurgAtlas enable training surgical foundation models with broad procedural coverage: we finetune Qwen3-VL-8B through a two-stage captioning-then-instruction pipeline and achieve competitive or state-of-the-art results on multiple established surgical benchmarks, including phase recognition, triplet detection, and reasoning question answering. More broadly, SurgAtlas provides a large native public video corpus that can support future large-scale pretraining of multimodal surgical AI systems and contribute to the development of next-generation foundation models for surgery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。