首个面向腹腔镜手术机器人的视觉语言动作模型评测基准
SurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical Robotics

- 构建从基础操作到完整术式分层的任务体系
- 自回归模型语义理解强,流匹配模型精度高但泛化弱
- 适合手术机器人、具身智能研究者参考
视觉-语言-动作(VLA)模型为手术机器人中的具身智能提供了新方向。尽管通用机器人领域已有多种VLA评测基准,但针对外科场景的标准化评估平台仍属空白。为此,我们提出SurgVLA-Bench,首个面向腹腔镜手术机器人任务的综合性评测基准。基于SurRoL仿真平台,构建了从原子动作到完整手术流程的分层任务体系,并设计多维度评估框架,涵盖动作准确率与语义一致性。我们系统评估了两类代表性模型:如OpenVLA的自回归模型,以及$π_{0}$、$π_{0.5}$和SmolVLA等流匹配模型。实验表明,自回归模型在语义理解方面表现更优,而流匹配模型通常具备更高任务精度,但存在泛化能力下降的问题。然而,即便最优模型仍远未达标,受限于内窥镜视野狭窄、视角受限及频繁遮挡等物理瓶颈,性能提升面临根本性挑战。代码与数据已开源。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models represent a promising direction for embodied intelligence in surgical robotics. Despite the prevalence of VLA benchmarks for general robotics, standardized evaluation platforms specifically designed for surgical contexts remain absent. To address this limitation, we present SurgVLA-Bench, the first comprehensive benchmark for evaluating VLA models in laparoscopic surgical robotics. Leveraging the SurRoL simulation platform, we construct a hierarchical task taxonomy ranging from atomic actions to complete surgical procedures, complemented by a multi-dimensional evaluation framework assessing action accuracy and semantic consistency. We then systematically evaluate two representative paradigms, including autoregressive models such as OpenVLA, and flow matching models such as $π_{0}$, $π_{0.5}$, and SmolVLA. Our experiments show that autoregressive models tend to excel in semantic understanding, while flow matching models often achieve higher task precision but may face generalization trade-offs. However, even the best-performing models remain far from satisfactory, as the constrained endoscopic field of view, restricted viewing angles, and frequent occlusions persist as fundamental physical bottlenecks. The code and data are available at https://github.com/VCL-HNU/SurgVLA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。