为机器人学习设计可衡量进展的奖励模型,解决传统成功信号无法反馈过程的问题。
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey

- 从输入输出接口、构建机制到评估基准,建立统一的进展奖励框架。
- 揭示不同方法在进展估计与奖励生成中的假设差异和实现原理。
- 适合研究机器人强化学习、进展奖励机制的学者参考。
机器人学习在动态环境和庞大的行为空间中进行。仅靠任务完成的最终成功信号无法判断当前行为是否在推进、停滞或倒退。因此,近期研究越来越多地探索进展奖励,以在任务执行过程中提供反馈。然而,现有文献缺乏统一框架:方法使用不同的观测信息、目标设定、输出信号、监督来源和评估协议,难以比较且结果验证不清。本文提出一个统一视角,将进展奖励建模分为三个连贯步骤:首先分析模型的外部接口,明确其接收的信息与输出的进展信号形式;其次剖析模型内部构建机制,揭示进展估计与奖励生成背后的假设与原理;最后考察支持方法的数据与基准,说明进展监督如何获取以及评估实际测量什么。三者共同连接了进展模型的本质、构建方式与质量验证。我们还总结了现有方法的主要局限,并讨论未来研究方向。
原文摘要 · Abstract (English)
Robotic learning takes place in dynamic environments with large behavior spaces. A terminal success signal only tells the robot whether the task is completed. It does not explain whether the current behavior is making progress, remaining unchanged, or undoing earlier progress. For this reason, recent studies have increasingly explored progress rewards that provide feedback during task execution. However, the current literature lacks a shared framework. Existing methods use different observations, goal specifications, output signals, supervision sources, and evaluation protocols. This makes it difficult to compare them and understand what their results actually validate. In this survey, we provide a unified view of progress reward modeling for robotic learning. We organize the field in three connected steps. We first study the interface of a progress model. This defines the problem from the outside by asking what information the model receives and what form of progress signal it produces. We then move inside the model and study the methods used to construct this signal. This reveals the different assumptions and mechanisms behind progress estimation and reward generation. Finally, we examine the data and benchmarks that support these methods. This shows how progress supervision is obtained and what different evaluations actually measure. Together, these three perspectives connect what a progress model is, how it is built, and how its quality is validated. We further summarize the main limitations of current approaches and discuss future research directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。