arXiv:2606.21572cs.RO2026-06被引 1

用细粒度视觉差异训练机器人批评者,提升操作成功率。

Robot Critics that Sweat the Small Stuff

论文配图:Robot Critics that Sweat the Small Stuff
图 1 · 摘自论文原文
  • 通过成功与失败轨迹构建成对监督信号微调视觉语言模型
  • 在真实世界任务中使策略成功率提升11%,仿真中提升5.9%
  • 能精准识别微小动作差异,适合高精度机器人操控场景

大型视觉语言模型具备世界和物体交互的先验知识,可在推理阶段作为批评者引导机器人策略走向成功。然而,闭环机器人操作需要判断成功与失败间的细微视觉差异,这仍是当前视觉语言模型的挑战。本文提出一种方法,利用策略生成的成功与失败轨迹构造成对进展监督信号来微调批评者。微调后的批评者在细粒度进展推理和微小失败检测方面表现优异,超越已有进展推理基线。此外,我们使用动作条件视频模型预测多个候选动作的视觉效果,并证明该批评者可正确识别应执行的成功动作,使真实世界任务平均策略成功率提升11%,仿真任务提升5.9%。

原文摘要 · Abstract (English)

Large vision-language models contain several priors about the world and object interactions, making them useful critics during inference to steer robot policies towards success. However, closed-loop robot manipulation requires judging small visual differences between success and failure, which remains a challenge for current VLMs. We introduce a method to fine-tune critics by constructing pairwise progress supervision using success and failure rollouts obtained from a policy. Our fine-tuned critic excels at fine-grained progress reasoning and subtle failure detection, outperforming prior progress reasoning baselines. Additionally, we use an action-conditioned video model to predict the visual effect of several candidate actions sampled from a policy, and show that our critic can correctly identify successful candidates to execute, improving the average policy success rate by 11% across real-world tasks and 5.9% across simulation tasks.

机器人视觉语言模型强化学习细粒度判断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。