arXiv:2503.02077cs.MAcs.AI2025-03ICML被引 14

用多阶段混合质量人类反馈提升多智能体强化学习的协作效果

M3HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed Quality

  • 通过多轮人类反馈迭代优化智能体策略,融合专家与非专家意见
  • 在复杂环境中显著优于现有方法,实现更稳健的多智能体协作
  • 适合希望低成本参与AI训练、重视可解释性的研究者

在多智能体强化学习(MARL)中设计有效奖励函数是重大挑战,常导致复杂协调环境中的次优或错位行为。我们提出多阶段混合质量人类反馈框架(M³HF),将不同水平人类的迭代反馈融入训练过程。通过在训练中暂停学习进行人工评估,利用大语言模型解析反馈,并基于预设模板和自适应权重(结合权重衰减与性能调整)更新奖励函数。该方法能有效整合多层次质量的人类见解,提升多智能体协作的可解释性与鲁棒性。在多个挑战性环境中的实验证明,M³HF显著优于当前最优方法,有效应对了MARL中奖励设计的复杂性,并支持更广泛的人类参与。

原文摘要 · Abstract (English)

Designing effective reward functions in multi-agent reinforcement learning (MARL) is a significant challenge, often leading to suboptimal or misaligned behaviors in complex, coordinated environments. We introduce Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed Quality ($\text{M}^3\text{HF}$), a novel framework that integrates multi-phase human feedback of mixed quality into the MARL training process. By involving humans with diverse expertise levels to provide iterative guidance, $\text{M}^3\text{HF}$ leverages both expert and non-expert feedback to continuously refine agents' policies. During training, we strategically pause agent learning for human evaluation, parse feedback using large language models to assign it appropriately and update reward functions through predefined templates and adaptive weights by using weight decay and performance-based adjustments. Our approach enables the integration of nuanced human insights across various levels of quality, enhancing the interpretability and robustness of multi-agent cooperation. Empirical results in challenging environments demonstrate that $\text{M}^3\text{HF}$ significantly outperforms state-of-the-art methods, effectively addressing the complexities of reward design in MARL and enabling broader human participation in the training process.

多智能体人类反馈强化学习协作优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。