arXiv:2608.30378cs.ROcs.AI2026-08

让机器人从成功轨迹中学习动作,只在好行为时执行。

PAVE: Predictive Alignment and Value-Guided Evolution for World-Action Policies

论文配图:PAVE: Predictive Alignment and Value-Guided Evolution for World-Action Policies
图 1 · 摘自论文原文
  • 用多时标预测学习场景演化规律
  • 通过价值评估筛选高质量动作轨迹
  • 保留直接执行路径,适合真实部署

直接视觉-语言-动作策略能高效生成连续机器人动作,但标准行为克隆存在两个互补缺陷:其表示未显式要求描述场景在多时间尺度上的演变,且部署轨迹质量不一却常被重复使用,未能区分有效动态与不良行为。我们提出 extit{PAVE},一种结合结果无关预测学习与结果相关策略优化的直接世界-动作策略。 extit{PAVE} 首先保留局部固定偏移的 JEPA 目标,并在剩余回合的 25%、50%、75% 和 100% 处加入轨迹相对的多时标状态转移对齐。这些仅训练目标要求当前策略表示同时保持局部物理变化与长程任务进展,无需向动作头提供显式未来标记。随后, extit{PAVE} 在累积部署轨迹上独立训练分布值评论器,计算与动作块对齐的 $N$-步优势,并将其转化为正、负或空文本条件用于流匹配动作器。因此,每个有效轨迹都能传递物理发生过程,而动作器仅在关联较优动作的条件下部署。多时标预测器与评论器在在线执行中被移除,保留了从当前观测、语言指令和本体感知直接生成动作的能力。在三个仿真基准上, extit{PAVE} 实现最强整体性能,同时保持直接动作器的在线执行路径。

原文摘要 · Abstract (English)

Direct vision-language-action policies generate continuous robot actions efficiently, but standard behavior cloning leaves two complementary gaps: their representations are not explicitly required to describe how the scene evolves over multiple time scales, and deployment trajectories of unequal quality are often reused without separating useful dynamics from undesirable behavior. We introduce \method, a direct world-action policy that combines outcome-agnostic predictive learning with outcome-aware policy improvement. \method first retains a local fixed-offset JEPA objective and adds trajectory-relative multi-horizon transition alignment at 25%, 50%, 75%, and 100% of the remaining episode. These training-only targets require the current policy representation to preserve both local physical changes and longer-range task progress, without supplying explicit future tokens to the action head. \method then trains an independent distributional value critic on cumulative deployment trajectories, computes action-chunk-aligned $N$-step advantages, and converts them into positive, negative, or null text conditions for a flow-matching actor. Thus, every valid trajectory can teach what physically happened, while the actor is deployed only under the condition associated with relatively better actions. The multi-horizon predictor and critic are removed from online execution, preserving direct action generation from the current observation, language instruction, and proprioception. \redclaim{Across the three simulation benchmarks, \method achieves the strongest overall performance while preserving the direct actor's online execution path.}

机器人控制强化学习多模态动作策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。