用正向示范学习机器人动作的正确变化,生成密集奖励。
Dream2Reward: Transition-Alignment Reward Models from Positive Demonstrations for Robotic Manipulation

- 从成功示范中学习理想动作变化的潜在空间,实现逐步对比
- 在错误转向、越界或停滞时仍能有效惩罚,保持奖励敏感性
- 无需失败标注或合成负样本,适合真实机器人在线学习
机器人策略学习需要在行为偏离成功示范时仍具信息量的密集奖励。基于进展的奖励虽可衡量观察是否沿预期轨迹推进,但在错误转移后仍可能保持高位。我们提出 Dream2Reward,通过正向示范学习语言条件下的成功潜在转移场。给定视觉历史至转移起点,模型预测成功执行的潜在位移,并通过符号方向与对称幅度一致性评估实际位移。该逐级比较机制在错误方向、越界和停滞等情况下仍能有效惩罚,即使结果看似有进展。Dream2Reward 不需失败标注、进展标签或合成负样本,生成稠密因果奖励。在机制诊断与共享轨迹评估中,其成功-失败分离更强,对低质量行为反馈更丰富。在在线与离线策略学习中,同一冻结奖励模型减少奖励黑客行为,提升下游性能,包括真实机器人操作。结果表明,将实际运动与预测成功变化对比,是将正向示范转化为密集奖励的有效方法。
原文摘要 · Abstract (English)
Learning robotic policies requires dense rewards that remain informative when behavior departs from successful demonstrations. Progress-based rewards estimate how far an observation has advanced along a nominal successful trajectory, but may remain high after an incorrect transition. We introduce Dream2Reward, which learns a language-conditioned successful latent transition field from positive demonstrations. Given the visual history up to a transition start, the model predicts the latent displacement associated with successful execution and scores the observed displacement through signed directional and symmetric magnitude agreement. This transition-level comparison penalizes wrong-direction, overshooting, and stagnant motion even when the resulting observation appears to show progress. Dream2Reward requires no failure annotations, progress labels, or synthetic negatives, and produces a dense causal reward. Across mechanism diagnostics and shared-trajectory evaluations, it provides stronger success-failure separation and more informative feedback on low-quality behavior than progress-based alternatives. Across online and offline policy learning, the same frozen reward model reduces reward hacking and supports stronger downstream performance, including in real-robot manipulation. These results show that comparing realized motion with predicted successful change provides an effective way to convert positive demonstrations into dense rewards for robot learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。