arXiv:2604.08168cs.ROcs.AI2026-04被引 11

用视频生成模型预测未来动作和价值,提升机器人长任务决策能力

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning

  • 用预训练视频生成器联合预测未来本体感知和价值信号
  • 在3个任务上实现最高80%成功率,准确追踪任务进展并发现错误
  • 适合需要长期规划与环境交互的机器人强化学习场景

视觉-语言-动作(VLA)模型通过大规模预训练推动了机器人操作发展,但现实部署仍受制于观测不完全和反馈延迟。强化学习通过价值函数评估任务进展并指导策略优化,但现有基于视觉-语言模型(VLM)的价值模型难以捕捉时序动态和物理交互,影响长周期任务中的可靠价值估计。本文提出ViVa,一种视频生成式价值模型,将预训练视频生成器重用于联合预测未来本体感知与标量价值。通过将价值估计锚定在预期的身体动态上,ViVa利用时空先验,使价值与前瞻能力内在耦合,超越静态快照。ViVa在三个任务的度量评估中达到当前最优表现,生成可靠的值信号,能精准追踪任务进展并检测执行错误。集成至RECAP后,平均成功率达80%,凸显视频生成模型在价值估计中的潜力。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models have advanced robot manipulation through large-scale pretraining, but real-world deployment remains challenging due to partial observability and delayed feedback. Reinforcement learning addresses this via value functions, which assess task progress and guide policy improvement. However, existing value models built on vision-language models (VLMs) struggle to capture temporal dynamics and physical interactions, undermining reliable value estimation in long-horizon tasks. In this paper, we propose ViVa, a video-generative value model that repurposes a pretrained video generator to jointly predict future proprioception and a scalar value. By grounding value estimation in anticipated embodiment dynamics, ViVa leverages spatiotemporal priors to intrinsically couple value with foresight beyond static snapshots. ViVa achieves state-of-the-art results in metric-based evaluation across three tasks, producing reliable value signals that accurately track task progress and detect execution errors. Integrated into RECAP, it achieves an average success rate of 80%, highlighting the promise of video-generative models for value estimation.

机器人学习价值模型视频生成强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。