用强化学习优化视频模型的视觉动态,提升机器人操控精度。
VAMPO: Policy Optimization for Improving Visual Dynamics in Video Action Models
- 将多步去噪建模为序列决策问题,通过奖励优化潜空间动态。
- 在模拟和真实任务中显著改善物体姿态与接触时机的准确性。
- 适合关注视频生成与机器人控制结合的研究者。
视频动作模型是视觉-语言-动作系统的重要基础,能从大规模视频数据中学习视觉动态并迁移到下游机器人控制。然而,当前基于扩散的视频预测器采用似然代理目标训练,虽能生成全局合理的结果,却未显式优化对操作至关重要的视觉动态精度。这种目标不匹配常导致物体位姿、空间关系和接触时机等细微误差,被下游策略放大。本文提出VAMPO,一种后训练框架,通过策略优化直接提升视频动作模型的视觉动态。核心思想是将多步去噪视为序列决策过程,以专家级潜空间视觉动态定义奖励。为实现高效优化,引入欧拉混合采样器,仅在首步注入随机性,确保低方差策略梯度估计的同时保持后续去噪轨迹连贯性。结合GRPO与可验证的非对抗性奖励,VAMPO在多样化的模拟与真实世界操作任务中均提升了任务相关的视觉动态,带来更优的下游动作生成与更强泛化能力。
原文摘要 · Abstract (English)
Video action models are an appealing foundation for Vision--Language--Action systems because they can learn visual dynamics from large-scale video data and transfer this knowledge to downstream robot control. Yet current diffusion-based video predictors are trained with likelihood-surrogate objectives, which encourage globally plausible predictions without explicitly optimizing the precision-critical visual dynamics needed for manipulation. This objective mismatch often leads to subtle errors in object pose, spatial relations, and contact timing that can be amplified by downstream policies. We propose VAMPO, a post-training framework that directly improves visual dynamics in video action models through policy optimization. Our key idea is to formulate multi-step denoising as a sequential decision process and optimize the denoising policy with rewards defined over expert visual dynamics in latent space. To make this optimization practical, we introduce an Euler Hybrid sampler that injects stochasticity only at the first denoising step, enabling tractable low-variance policy-gradient estimation while preserving the coherence of the remaining denoising trajectory. We further combine this design with GRPO and a verifiable non-adversarial reward. Across diverse simulated and real-world manipulation tasks, VAMPO improves task-relevant visual dynamics, leading to better downstream action generation and stronger generalization. The homepage is https://vampo-robot.github.io/VAMPO/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。