arXiv:2604.03037cs.ROcs.AI2026-04被引 10

用相对优势替代绝对进度,实现低成本高精度的长程机械臂操作奖励建模。

ARM: Advantage Reward Modeling for Long-Horizon Manipulation

论文配图:ARM: Advantage Reward Modeling for Long-Horizon Manipulation
图 1 · 摘自论文原文
  • 通过进展、倒退、停滞三态标注,降低人工标注负担。
  • 在毛巾折叠任务中达99.4%成功率,数据效率显著优于现有视觉语言模型基线。
  • 适合需要少人工干预的长序列机器人操作场景。

长程机器人操作对强化学习仍具挑战性,因其稀疏奖励难以提供有效信用分配。实际策略改进依赖于更丰富的中间监督,如密集进度奖励,但此类标注成本高,且不适用于回溯或恢复等非单调行为。为此,我们提出优势奖励建模(ARM),将难以量化的绝对进度转化为相对优势估计。引入一种低成本的三态标注策略——进展、倒退、停滞,减少人类认知负担并保证跨标注者一致性。利用这些直观信号训练后,ARM可自动为完整示范和碎片化DAgger式数据生成进度标注。将其集成至离线强化学习流程中,可实现自适应动作-奖励重加权,有效过滤低质量样本。该方法在具有挑战性的长程毛巾折叠任务中取得99.4%的成功率,相较于当前VLA基线展现出更高的稳定性和数据效率,且政策训练期间近乎零人工干预。

原文摘要 · Abstract (English)

Long-horizon robotic manipulation remains challenging for reinforcement learning (RL) because sparse rewards provide limited guidance for credit assignment. Practical policy improvement thus relies on richer intermediate supervision, such as dense progress rewards, which are costly to obtain and ill-suited to non-monotonic behaviors such as backtracking and recovery. To address this, we propose Advantage Reward Modeling (ARM), a framework that shifts from hard-to-quantify absolute progress to estimating relative advantage. We introduce a cost-effective tri-state labeling strategy -- Progressive, Regressive, and Stagnant -- that reduces human cognitive overhead while ensuring high cross-annotator consistency. By training on these intuitive signals, ARM enables automated progress annotation for both complete demonstrations and fragmented DAgger-style data. Integrating ARM into an offline RL pipeline allows for adaptive action-reward reweighting, effectively filtering suboptimal samples. Our approach achieves a 99.4% success rate on a challenging long-horizon towel-folding task, demonstrating improved stability and data efficiency over current VLA baselines with near-zero human intervention during policy training.

机器人操作强化学习奖励建模数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。