arXiv:2605.19580cs.RO2026-05

让机器人更靠谱:识别关键决策动作并重点优化

PAPO-VLA: Planning-Aware Policy Optimization for Vision-Language-Action Models

论文配图:PAPO-VLA: Planning-Aware Policy Optimization for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 通过动作变化与任务结果联合判断,识别出关键决策动作
  • 利用因果充分性与必要性量化决策动作的重要性
  • 在优化中强化重要决策动作,提升任务成功率

视觉-语言-动作(VLA)模型在语言引导的机器人任务中展现出巨大潜力。然而,由于操作任务依赖闭环交互,每一步动作都会影响后续执行,因此确保VLA策略的可靠性仍具挑战。本文重新审视执行过程中的VLA策略,认为其兼具规划者与执行者双重角色:规划者做出任务导向的决策以改变执行方向,执行者则通过密集连续动作实现这些决策。这一视角表明,提升VLA可靠性需特别关注规划动作。现有优化方法虽可模仿动作或改进完整轨迹,但通常未显式识别规划动作,也未衡量其对任务成功的重要性。为此,本文提出规划感知的策略优化方法(PAPO-VLA):首先通过动作变异与轨迹结果联合判断识别规划动作;其次利用因果充分性与必要性估计其重要性;最后将该重要性融入GRPO的优势估计中。如此,重要规划动作获得更强优化聚焦,同时整个轨迹仍受轨迹级反馈指导。多基准测试实验验证了PAPO-VLA的有效性。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models show promising ability in language-guided robotic tasks. However, making VLA policies reliable remains challenging, because a manipulation task is completed through closed-loop interaction, where each action affects subsequent execution. To analyze this problem, we revisit VLA policy during execution and argue that a VLA policy acts both as a planner, which makes task-oriented decisions that change the direction of execution, and as an executor, which realizes these decisions through dense continuous actions. This view suggests that improving VLA reliability requires particular attention to planning actions. Existing optimization methods can imitate actions or improve complete trajectories, but they usually do not explicitly identify planning actions or measure their importance for task success. To address this issue, we propose Planning-Aware Policy Optimization for VLA models (PAPO-VLA). PAPO-VLA first identifies planning actions by jointly considering action variation and trajectory outcome, then estimates their importance through causal sufficiency and causal necessity, and finally incorporates this importance into GRPO advantage estimation. In this way, more important planning actions receive stronger optimization emphasis, while the whole trajectory is still optimized by trajectory-level feedback. Experiments on multiple benchmarks demonstrate the effectiveness of PAPO-VLA.

机器人强化学习决策优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。