提出统一框架,分析动态奖励设计如何影响强化学习效果
A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning
- 区分参数更新与状态依赖变化,拆解奖励塑造机制
- 12类方法对比发现:适应性奖励在多数场景下仍可保最优策略
- 适合研究动态奖励设计或深度强化学习稳定性的学者
稀疏、延迟且信息量弱的奖励仍是强化学习效率的核心障碍。奖励塑造通过引入辅助信号补充任务奖励,可加速学习,经典理论保证:当辅助项为时间不变势能的折现差时,最优策略不变。然而现代强化学习中,学习者和可用指导信息随训练动态演变:价值估计提升、新颖性降低、反馈变化、预测模型优化。自适应奖励机制广泛存在于探索、贝叶斯推断、人机协同、自动奖励设计及基础模型方法中。本文提出统一分析框架,比较动态奖励塑造与邻近自适应机制。框架区分参数修订与状态依赖变化,分离加性塑造、奖励替换与奖励邻近引导,并按时间、信息与理论维度组织现有方法。基于此框架,对12类方法进行对比分析,揭示了在当代深度强化学习流水线、回放缓冲区、自举评论器和奖励归一化条件下,最优性保证仍成立的条件,同时暴露了适应速率与学习器稳定性之间未解决的关系。
原文摘要 · Abstract (English)
Sparse, delayed, and weakly informative rewards remain central obstacles to efficient reinforcement learning. Reward shaping addresses these limitations by supplementing the task reward with an auxiliary signal that can accelerate learning while, in the classical setting, the original objective remains the evaluation criterion. Established theory guarantees safety for fixed shaping signals: potential-based reward shaping preserves optimal policies when the auxiliary term is the discounted difference of a time-invariant potential. In contemporary reinforcement learning systems, however, both the learner and the information available for guidance evolve during training: value estimates improve, novelty diminishes, feedback shifts, and predictive models are refined. Adaptive reward mechanisms occur across exploration, Bayesian inference, human-in-the-loop learning, automated reward design, and foundation-model-based approaches. This study introduces a unified analytical framework for comparing dynamic reward shaping and neighbouring adaptive reward mechanisms. The proposed framework distinguishes parametric revision from state-dependent variation, separates additive shaping from reward replacement and reward-adjacent guidance, and organises existing methods along temporal, informational, and theoretical dimensions. Using this framework, twelve method families are comparatively analysed. The framework further highlights the conditions under which optimality guarantees survive contemporary deep reinforcement learning pipelines, replay buffers, bootstrapped critics, and reward normalisation, while exposing the unresolved relationship between adaptation rate and learner stability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。