arXiv:2606.31377cs.ROcs.AI2026-06

用专家视频自动生成分阶段的密集奖励,让机器人更高效学会复杂操作。

Stage-Transition Dense Reward Modeling for Reinforcement Learning

论文配图:Stage-Transition Dense Reward Modeling for Reinforcement Learning
图 1 · 摘自论文原文
  • 从专家演示中解析任务阶段结构,生成阶段转移与阶段内进度双重奖励信号。
  • 在14个任务上提升采样效率和成功率,部分任务超越人工设计奖励。
  • 支持真实机器人部署,对视觉噪声鲁棒,奖励分配更准确合理。

长时程机器人操作中的强化学习常受限于稀疏延迟的奖励,而手动设计密集奖励信号成本高且易受环境变化影响。本文提出阶段转移密集奖励(STDR)框架,将未结构化的专家视频转化为逻辑自洽的密集奖励,用于从零训练强化学习智能体。STDR利用语义理解从示范中推断任务阶段结构,在在线训练中提供两类互补信号:(i) 阶段转移反馈,实现目标导向奖励;(ii) 阶段内进度反馈,提供精细步骤指引。此外,集成分布外检测机制与抓取调控模块,增强鲁棒性并防止奖励欺骗。在MetaWorld、ManiSkill和Franka Kitchen上的14项操作任务实验表明,STDR在多个基线之上持续提升样本效率与成功率达显著,部分任务性能媲美甚至超越人工设计的密集奖励。真实机器人测试进一步显示,成功执行时奖励稳定且与进展对齐,失败时奖励显著降低,表明其对视觉噪声具有强鲁棒性,并在不同场景下具备更精准的奖励校准能力。

原文摘要 · Abstract (English)

Reinforcement learning for long-horizon robotic manipulation is often limited by sparse and delayed rewards, while manually designing dense shaping signals is costly and brittle to changes in environments and object configurations. This work proposes Stage-Transition Dense Reward (STDR), a visual reward-learning framework that converts unstructured expert videos into logically grounded dense rewards for training RL agents from scratch. STDR leverages semantic understanding to infer a task's stage structure from demonstrations, and delivers two complementary learning signals during online training: (i) stage-transition feedback that provides goal-directed reward, and (ii) within-stage progress feedback that supplies fine-grained guidance toward completing each stage. Furthermore, an out-of-distribution (OOD) detection mechanism and a grasping regulation module are integrated to enhance robustness and prevent reward hacking. Experiments on 14 manipulation tasks across MetaWorld, ManiSkill, and Franka Kitchen show that STDR consistently improves sample efficiency and success rates over multiple baselines, and matches or surpasses handcrafted dense rewards on several challenging tasks. Real-robot evaluations further indicate that STDR assigns stable, progress-aligned rewards on successful executions while producing appropriately low rewards for failures, suggesting robustness to visual noise and better-calibrated reward assignment across settings.

强化学习机器人操作密集奖励阶段建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。