用最简结构实现专家演示的奖励分配,效果不输复杂方法。
Minimal Ingredients for Reward Assignment from Expert Demonstrations
- 仅靠轨迹相似度就能有效指导离线强化学习。
- 轻量级时间对齐能提升在线学习表现,多示范时更关键。
- 适合追求简洁高效的模仿学习研究者参考。
从稀缺示范中进行奖励分配是离线与在线模仿学习的核心挑战。常见策略是根据学习者轨迹与专家示范的接近程度分配奖励。尽管这一原则支撑了众多现有方法,其决定性能的关键要素仍缺乏系统研究。本文从邻近性近似与时间对齐两个设计维度出发,在32个涵盖离线与在线场景的基准上,结合三种下游强化学习算法进行实证分析。结果表明:(1) 离线场景中,仅依赖邻近性即可捕获有效奖励结构;(2) 轻量级时间对应在离线中带来微小增益,但在在线或存在多个示范时至关重要。此外,我们补充了一项轻量理论,阐明简单邻近近似何时足够。整体表明,在模仿学习中应优先采用算法极简主义,避免过早引入复杂奖励设计。
原文摘要 · Abstract (English)
Reward assignment from scarce demonstrations is a key challenge in both offline and online imitation learning. A common and intuitive strategy assigns rewards according to how closely learner trajectories match expert demonstrations. Although this principle underlies many existing methods, the core ingredients that drive performance remain systematically underexplored. We therefore ask: what is the minimal structure that reward assignment must encode to achieve effective downstream RL performance across settings? We approach this question along two design axes: proximity approximation and temporal alignment. Across 32 benchmarks spanning offline and online settings, and with three downstream RL algorithms, our empirical findings suggest: (1) In offline regimes, proximity alone captures the reward structure necessary for effective offline RL, while (2) lightweight temporal correspondence provides consistent gains that are modest offline but essential online or in the presence of multiple demonstrations. We further complement our offline results with a lightweight theory characterizing when simple proximity approximation suffices. Overall, these findings advocate algorithmic minimalism in reward design before introducing complex schemes in both offline and online imitation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。