arXiv:2609.07211cs.RO2026-09

用视觉语言模型追踪手术操作进度,让机器人从失败中学习局部成功经验。

Phase-and-First-Arrival VLM Feedback for Sparse-Reward Reinforcement Learning in Surgical Manipulation

论文配图:Phase-and-First-Arrival VLM Feedback for Sparse-Reward Reinforcement Learning in Surgical Manipulation
图 1 · 摘自论文原文
  • 每条轨迹只问一次VLM,识别出最早达到的可验证任务阶段
  • 模拟实验中成功率提升至75.2%,高于基线的52.1%
  • 适用于复杂手术操作中的部分行为复用与实时反馈

稀疏奖励反馈限制了机器人从复杂操作失败中学习的能力。未成功的多阶段手术尝试中可能包含值得复用的抓取、提举或转移动作。在稀疏奖励强化学习中,终端奖励将这些尝试统一归为相同结果,而标量视觉语言模型(VLM)评分既无法判断哪些进展应获奖励,也无法确定其发生时间。我们提出相位与首次到达反馈:每个记录的实验仅需一次VLM查询,即可识别出最早被视觉验证的任务阶段及其首次到达时刻,使学习者能复用部分行为并准确定位奖励时间。我们在SurgPhaseBench上实现该方法,该基准涵盖刚性与柔性任务的结构化阶段。在五个模拟任务中,我们的方法平均成功率达75.2%,优于基于对比语言图像预训练(CLIP)的基线52.1%;当仅改变反馈表示时,优势依然存在。在硬件上,同一记录支持自主拾块与滑动恢复。结果表明,轨迹级视觉监督可在稀疏奖励控制中保留部分进展并提供必要的时间信用。

原文摘要 · Abstract (English)

Sparse outcome feedback limits what robots can learn from unsuccessful attempts at complex manipulation. Failed multi-stage surgical attempts can contain grasps, lifts, or transfers worth reusing. In sparse-reward reinforcement learning, terminal rewards collapse such attempts to the same outcome, while scalar vision-language model (VLM) ratings reveal neither what progress merits credit nor when it occurred. We introduce phase-and-first-arrival feedback: one VLM query per recorded episode identifies the furthest visually verified task phase and when that phase is first reached, allowing the learner to reuse partial behavior and localize credit. We instantiate it in SurgPhaseBench, a phase-structured suite spanning rigid and deformable tasks, and evaluate it in simulation and hardware. Across five simulated tasks, our method reaches 75.2% mean success, compared with 52.1% for a reward based on Contrastive Language-Image Pre-training (CLIP) using the same visual input; the advantage persists when only the feedback representation changes. On hardware, the same record supports autonomous block picking and slip recovery. Together, these results show that trajectory-level visual supervision can preserve partial progress while providing the temporal credit needed for sparse-reward control.

强化学习手术机器人视觉语言模型稀疏奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。