arXiv:2507.18867cs.LGcs.AI2025-07被引 4

让智能体学习人类专家偏好,自动生成高效探索的个体奖励

Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise

  • 用人类专家偏好指导每个智能体的动作分布
  • 在稀疏奖励环境下提升团队性能,优于主流基线
  • 适合需要高效探索的复杂多智能体任务

多智能体强化学习(MARL)在仅有团队奖励且奖励稀疏的环境中面临高效探索难题。现有方法依赖人工设计的奖励函数,缺乏高层次智能,泛化能力差。本文提出LIGHT框架,通过端到端方式将人类专家知识融入MARL。该方法结合个体动作分布与人类偏好分布,设计基于Q-learning相关可操作表征变换的内在奖励,使智能体动作偏好对齐人类专家,同时最大化联合动作价值。实验表明,LIGHT在多个挑战性场景中性能优于主流基线,且在不同稀疏奖励任务间具备更强的知识复用能力。

原文摘要 · Abstract (English)

Efficient exploration in multi-agent reinforcement learning (MARL) is a challenging problem when receiving only a team reward, especially in environments with sparse rewards. A powerful method to mitigate this issue involves crafting dense individual rewards to guide the agents toward efficient exploration. However, individual rewards generally rely on manually engineered shaping-reward functions that lack high-order intelligence, thus it behaves ineffectively than humans regarding learning and generalization in complex problems. To tackle these issues, we combine the above two paradigms and propose a novel framework, LIGHT (Learning Individual Intrinsic reward via Incorporating Generalized Human experTise), which can integrate human knowledge into MARL algorithms in an end-to-end manner. LIGHT guides each agent to avoid unnecessary exploration by considering both individual action distribution and human expertise preference distribution. Then, LIGHT designs individual intrinsic rewards for each agent based on actionable representational transformation relevant to Q-learning so that the agents align their action preferences with the human expertise while maximizing the joint action value. Experimental results demonstrate the superiority of our method over representative baselines regarding performance and better knowledge reusability across different sparse-reward tasks on challenging scenarios.

多智能体强化学习人类偏好稀疏奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。