arXiv:2606.18247cs.ROcs.AI2026-06被引 1

让机器人在运行时自我验证并优化策略,无需额外训练。

Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement

论文配图:Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement
图 1 · 摘自论文原文
  • 用视觉验证器在推理时评估动作,实现无需训练的策略调整。
  • 验证后的轨迹使离线微调后性能显著提升,媲美专家示范。
  • 部署中自动生成高质量数据,无需人工干预,适合实际应用。

真实世界中的机器人应能从经验中学习并持续改进,这需要具备实践与反馈学习的机制。本文提出VERITAS——一种通用机器人策略的生成-验证框架,用于推理时策略引导与自主优化。我们以预训练的通用机器人策略作为“生成器”,搭配一个无梯度的“视觉验证器”在推理阶段评估动作。该框架实现了推理时的策略调控,可提升性能而无需额外训练。实验表明,推理时验证优于未验证的通用策略,且无需额外示范数据。此外,经验证的轨迹为离线策略优化提供了有效监督:基于这些自生成轨迹微调的策略获得一致性能提升。值得注意的是,使用验证轨迹进行后训练的效率可媲美专家示范,且无需人工参与。结果表明,推理时验证是一种实用且可扩展的部署期策略优化机制。

原文摘要 · Abstract (English)

Robots deployed in the real world should learn from their experience and improve over time. This requires a mechanism of practicing and learning from feedback. In this paper, we propose VERITAS, a generator-verifier framework for generalist robot policies for inference-time policy steering and self-improvement. We use a pre-trained generalist robot policy as a ``generator'' and pair it with a gradient-free ``visual verifier'' that evaluates actions at inference time. This framework enables inference-time steering that improves policy performance without additional training. We demonstrate that inference-time verification consistently outperforms vanilla generalists without training on additional demonstration data. Additionally, we demonstrate that the verified rollouts provide effective supervision for offline policy improvement: policies fine-tuned on verified self-generated trajectories achieve consistent performance gains. Notably, we find that post-training with verified rollouts achieves comparable efficiency to expert demonstrations, while requiring no human interventions. Our results highlight inference-time verification as a practical and scalable mechanism for improving robotic policies during deployment.

机器人策略优化自学习视觉验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。