构建首个机器人手术通用模型,统一训练与评估标准。
OmniRAS: Standardizing Foundation Model Training and Evaluation in Robot-Assisted Surgery
- 基于10万+小时手术视频,训练10亿和20亿参数视觉编码器。
- 在6项任务中超越现有模型,冻结权重下仍表现最优。
- 提供标注数据集与评估协议,推动手术AI标准化发展。
针对机器人手术领域缺乏基础模型的问题,本文提出OmniRAS,一个包含10亿和20亿参数的V-JEPA-2.1视觉编码器系列,并详细阐述其训练过程。首先,发布两个密集标注的机器人胆囊切除数据集:OmniRAS-PR 和多标签YT-Chole工具-动作-目标任务数据集,首次采用三元组标注方式,配套划分方案、探测协议及评分者间一致性研究,验证共享阶段本体的可靠性。其次,利用最多256个计算节点,以全局批量6,144,在总计约2,650小时的手术视频(其中51%为机器人手术)上完成持续预训练,分析算力与数据构成。第三,分别在冻结编码器与最终四层微调两种设置下,对六项任务(包括三元组、阶段、步骤识别、动作分割与检测)进行评估,共完成254次下游实验,其中109次为部分主干微调。最佳的OmniRAS模型在所有任务类别中均取得最强适配结果,而冻结权重下的差异较小。
原文摘要 · Abstract (English)
Few foundation models exist for robot-assisted surgery, partly because large robotic-surgery video corpora are difficult to assemble and existing models are evaluated mostly on laparoscopic benchmarks. Further, most existing models are evaluated on a small set of public benchmarks, mostly focused on laparoscopic surgery. We present OmniRAS, a family of 1B- and 2B-parameter V-JEPA-2.1 encoders for robot-assisted surgery, and detail their training. First, we release two densely annotated robotic-cholecystectomy datasets: OmniRAS-PR and a multi-label YT-Chole tool-verb-target task, the first triplet-style annotation for robotic cholecystectomy, together with splits, probe protocols, and an inter-rater study validating the shared phase ontology. Second, we document continued pretraining at up to 256 compute nodes with global batch 6,144 over 19 sources totaling approximately 2,650 hours of surgical video, 51% robotic, and analyze compute and data composition. Third, we evaluate against raw V-JEPA-2.1 and specialized surgical models on six tasks spanning triplet, phase, and step recognition, action segmentation, and detection, under frozen-encoder and final-four-block fine-tuning regimes. Across three seeds, this yields 254 downstream runs, including 109 with partial backbone fine-tuning. The best OmniRAS models achieve the strongest adapted results across all task families, while frozen differences are smaller.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。