用分阶段奖励模型提升长时序机器人操作成功率
SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation
- 基于视频和语言标注,分阶段预测任务进展
- T恤折叠成功率达83%(平铺)和67%(皱缩),远超基线
- 适合复杂变形物体操作,减少对高质量演示依赖
大规模机器人学习在复杂操作任务上取得进展,但涉及可变形物体的长时序、高接触任务仍具挑战,主要因示范质量不一致。本文提出一种阶段感知的视频奖励建模框架,结合自然语言子任务标注,联合预测任务阶段与细粒度进展,生成跨时长示范的一致标签。该方法避免了基于帧索引标注的脆弱性,在如T恤折叠等任务中提供稳定监督。所提奖励模型对示范差异具有鲁棒性,可泛化至分布外场景,并有效提升下游策略训练效果。在此基础上,我们提出奖励对齐行为克隆(RA-BC),根据奖励估计筛选并重加权示范。实验表明,该方法在真实世界滚开展示和人工评估中均显著优于基线。在T恤折叠任务中,从平铺状态达到83%成功率,从皱缩状态达67%,而传统行为克隆仅分别为8%和0%。结果表明,奖励建模是长时序机器人操作中可扩展且注释高效的解决方案。
原文摘要 · Abstract (English)
Large-scale robot learning has made progress on complex manipulation tasks, yet long horizon, contact rich problems, especially those involving deformable objects, remain challenging due to inconsistent demonstration quality. We propose a stage-aware, video-based reward modeling framework that jointly predicts task stage and fine-grained progress, using natural language subtask annotations to derive consistent labels across variable-length demonstrations. This avoids the brittleness of frame index based labeling and provides stable supervision even in tasks like T-shirt folding. Our reward model is robust to demonstration variability, generalizes to out-of-distribution scenarios, and improves downstream policy training. Building on it, we introduce Reward-Aligned Behavior Cloning (RA-BC), which filters and reweights demonstrations based on reward estimates. Experiments show that our method significantly outperforms baselines in both real-world rollouts and human validation. On T-shirt folding, we achieve 83% success from the flattened state and 67% from the crumpled state, compared to 8% and 0% with vanilla BC. Overall, our results highlight reward modeling as a scalable and annotation-efficient solution for long horizon robotic manipulation. Project website: https://qianzhong-chen.github.io/sarm.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。