构建大规模手术解剖数据集并提出上下文感知识别模型,提升微创手术视觉理解能力。
Surgical Anatomy Recognition with Context Learning using Foundation Representations

- 利用基础模型嵌入与轻量时序推理融合手术上下文信息
- 在12万+帧数据上实现高精度时序一致的解剖结构分割
- 适合开发临床可用的微创手术辅助系统
准确识别解剖结构对安全有效的微创手术至关重要,但受限于标注数据少且现有方法多针对自然场景,该领域研究仍不充分。本文提出联合数据集与模型框架以推动微创手术中的解剖感知。首先,构建ATLAS-120k,一个包含超过12万帧的大型片段级语义分割数据集,覆盖100个手术视频、14种术式及多种模态(腹腔镜与机器人辅助手术),通过专家人工标注、自动传播、迭代优化与外科医生验证的可扩展流程生成高质量标注。其次,提出ATLAS(Anatomy Recognition with Context Learning using Foundation Representations)模型,专为手术解剖识别设计。不同于传统侧重目标跟踪的方法,ATLAS结合基础模型嵌入与轻量时序推理,引入术式类型、手术阶段和短期视觉记忆等上下文线索,实现时序一致且准确的预测,同时保持实时可行性。数据集与模型已开源,支持临床导向的微创手术引导系统开发。
原文摘要 · Abstract (English)
Accurate recognition of anatomical structures is essential for safe and effective minimally invasive surgery (MIS), yet it remains underexplored in surgical computer vision due to limited annotated data and methods tailored primarily to natural scenes. In this work, we present a combined dataset and model framework to advance anatomy-aware perception in MIS. First, we introduce ATLAS-120k, a large-scale clip-level semantic segmentation dataset comprising over 120,000 annotated frames from 100 surgical videos spanning 14 procedures and multiple modalities, including laparoscopic and robot-assisted surgery. The dataset captures substantial procedural variability and was created using a scalable annotation pipeline that integrates expert manual labeling, automated propagation, iterative refinement, and surgeon verification to ensure high-quality annotations. Second, we propose ATLAS (Anatomy Recognition with Context Learning using Foundation Representations), a video semantic segmentation model specifically designed for surgical anatomy recognition. Unlike conventional approaches that emphasize object tracking, ATLAS leverages foundation-model embeddings together with lightweight temporal reasoning to incorporate contextual cues such as procedure type, surgical phase, and short-term visual memory. This design enables temporally consistent and accurate predictions while maintaining real-time feasibility. Together, the dataset and model establish a practical foundation for robust surgical scene understanding and support the development of clinically applicable guidance systems for minimally invasive surgery. The models, dataset annotations and annotation platform are publicly available at: https://github.com/TimJaspers0801/ATLAS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。