arXiv:2503.21949cs.LG2025-03被引 4

设计智能奖励信号,让机器人更快学会正确行为。

Reward Design for Reinforcement Learning Agents

  • 从专家视角出发,设计可解释的奖励信号加速学习
  • 根据学习进度动态调整奖励,保持行为对齐
  • 让智能体自主设计奖励,实现自我优化

奖励函数在强化学习中至关重要,决定智能体如何决策。复杂任务需要精心设计的奖励,以有效引导学习并避免意外行为。本文探讨奖励信号的关键作用,分析其对智能体行为和学习动态的影响,解决延迟、模糊或复杂的奖励问题。研究分三部分:第一,基于专家策略与价值函数,设计能加速收敛的可解释奖励;第二,提出自适应可解释奖励设计方法,根据学习状态动态调整;第三,构建元学习框架,使智能体无需专家输入即可在线自设计奖励,形成自我改进的反馈循环。

原文摘要 · Abstract (English)

Reward functions are central in reinforcement learning (RL), guiding agents towards optimal decision-making. The complexity of RL tasks requires meticulously designed reward functions that effectively drive learning while avoiding unintended consequences. Effective reward design aims to provide signals that accelerate the agent's convergence to optimal behavior. Crafting rewards that align with task objectives, foster desired behaviors, and prevent undesirable actions is inherently challenging. This thesis delves into the critical role of reward signals in RL, highlighting their impact on the agent's behavior and learning dynamics and addressing challenges such as delayed, ambiguous, or intricate rewards. In this thesis work, we tackle different aspects of reward shaping. First, we address the problem of designing informative and interpretable reward signals from a teacher's/expert's perspective (teacher-driven). Here, the expert, equipped with the optimal policy and the corresponding value function, designs reward signals that expedite the agent's convergence to optimal behavior. Second, we build on this teacher-driven approach by introducing a novel method for adaptive interpretable reward design. In this scenario, the expert tailors the rewards based on the learner's current policy, ensuring alignment and optimal progression. Third, we propose a meta-learning approach, enabling the agent to self-design its reward signals online without expert input (agent-driven). This self-driven method considers the agent's learning and exploration to establish a self-improving feedback loop.

强化学习奖励设计智能体自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。