发现强化学习中关键步骤无奖励会导致学习失败,提出零激励动态新视角。
Zero-Incentive Dynamics: a look at reward sparsity through the lens of unrewarded subgoals
- 识别出关键子目标无直接奖励时的结构难题
- 顶尖算法在无即时奖励下性能显著下降
- 适合研究强化学习内在机制的学者参考
本文重新审视了奖励频率反映任务难度的普遍假设。我们识别并形式化了一种结构性挑战:当必要子目标无法直接获得奖励时,现有策略学习方法会失效。这类场景被称为零激励动态,即对成功至关重要的状态转移始终未被奖励。我们证明,当前最先进的深度子目标算法无法有效利用此类动态,且学习性能高度依赖子目标完成与最终奖励之间的时间接近度。这些发现揭示了现有方法的根本局限,并指向需要不依赖即时激励来推断潜在任务结构的新机制。
原文摘要 · Abstract (English)
This work re-examines the commonly held assumption that the frequency of rewards is a reliable measure of task difficulty in reinforcement learning. We identify and formalize a structural challenge that undermines the effectiveness of current policy learning methods: when essential subgoals do not directly yield rewards. We characterize such settings as exhibiting zero-incentive dynamics, where transitions critical to success remain unrewarded. We show that state-of-the-art deep subgoal-based algorithms fail to leverage these dynamics and that learning performance is highly sensitive to the temporal proximity between subgoal completion and eventual reward. These findings reveal a fundamental limitation in current approaches and point to the need for mechanisms that can infer latent task structure without relying on immediate incentives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。