arXiv:2608.15026cs.RO2026-08

让机器人长时操作任务中每一步都获得精准奖励,提升成功率。

PACE: Phase-Progress-Aware Credit for Long-Horizon Embodied Manipulation

论文配图:PACE: Phase-Progress-Aware Credit for Long-Horizon Embodied Manipulation
图 1 · 摘自论文原文
  • 通过分阶段进度感知的批评者,判断每一步是否推进任务
  • 在仿真和真实机械臂上显著优于最强基线,提升任务完成率
  • 适合长时序、多阶段的机器人操控任务优化

视觉-语言-动作(VLA)模型的后训练通常依赖专家示范和策略交互轨迹。然而,在长时程操作任务中,单个实验周期常包含数百个控制步骤和多个阶段,而成功或失败仅在任务结束时才可判定。因此,策略优化需要步骤级的信用分配信号来区分推动任务进展的行为与停滞或倒退行为。本文提出PACE框架,聚焦于阶段进度感知的批评者。该框架包含两个核心模块:(1) 全局-局部协同价值修正批评者(GLC-Critic),在局部时间窗口内聚合视觉与运动差异特征,推断每一步的阶段与阶段内进展,并对离散剩余成本分布进行残差修正,实现步骤级信用分配;(2) 逐步策略蒸馏(PPD),通过任务特定阈值将信用转化为正负条件,训练一个受信用条件驱动的动作生成策略:首先用高信用正样本保护预训练策略,再融合正负信用学习质量边界,并在推理时通过条件输出差异放大高信用行为。大量仿真实验和多样化的真实世界机械臂实验表明,PACE在各项指标上均显著优于最强基线。

原文摘要 · Abstract (English)

Post-training of vision-language-action (VLA) models typically relies on expert demonstrations and policy interaction trajectories. However, in long-horizon manipulation, a single episode often spans hundreds of control steps and multiple phases, while success or failure is only revealed at episode termination. Policy improvement therefore requires step-level credit signals to distinguish behaviors that advance the task from those that stall or regress. We present PACE, a credit-assignment framework for post-training on long-horizon manipulation, centered on a phase-progress-aware critic. PACE consists of two key modules: (1) the Global-Local Cooperative Value-Correction Critic (GLC-Critic) aggregates visual and motion-difference features within local temporal windows to infer the phase and intra-phase progress of each step, and applies residual correction to a discretized remaining-cost distribution accordingly, enabling step-level credit assignment; (2) Progressive Policy Distillation (PPD) converts credit into positive and negative conditions via task-wise thresholds and trains a credit-conditioned action generation policy: it first protects the pretrained policy with high-credit positive samples, then incorporates all positive and negative credits to learn the quality boundary, and at inference amplifies high-credit behaviors through the difference between conditional outputs. Extensive simulation experiments and diverse real-world robotic-arm experiments demonstrate that PACE consistently achieves significant improvements over the strongest baseline.

机器人操控信用分配长时任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。