arXiv:2503.03796cs.MAcs.AI2025-03被引 7

用人类隐式偏好优化无人艇集群强化学习策略

Human Implicit Preference-Based Policy Fine-tuning for Multi-Agent Reinforcement Learning in USV Swarm

  • 通过分层反馈机制区分个体、交互和团队层面的偏好
  • 在避障、区域约束等场景下提升集群行为一致性与公平性
  • 结合大模型评估,降低人工标注成本,适合智能集群系统研究者

多智能体强化学习(MARL)在无人水面艇(USV)集群执行搜救、监视和护航等任务中展现出潜力。然而,将系统行为与用户偏好对齐面临挑战,因专家直觉难以转化为奖励函数。为此,我们提出一种面向MARL的基于人类反馈的强化学习方法(RLHF),通过代理级反馈系统将反馈分为个体内部、智能体间和团队内三类,解决信用分配难题。为克服直接人工反馈的困难,采用大语言模型(LLM)评估器,在区域约束、碰撞规避和任务分配等反馈场景下验证方法有效性。实验表明,该方法能有效优化USV集群策略,同时保持行为公平性与性能一致性。

原文摘要 · Abstract (English)

Multi-Agent Reinforcement Learning (MARL) has shown promise in solving complex problems involving cooperation and competition among agents, such as an Unmanned Surface Vehicle (USV) swarm used in search and rescue, surveillance, and vessel protection. However, aligning system behavior with user preferences is challenging due to the difficulty of encoding expert intuition into reward functions. To address the issue, we propose a Reinforcement Learning with Human Feedback (RLHF) approach for MARL that resolves credit-assignment challenges through an Agent-Level Feedback system categorizing feedback into intra-agent, inter-agent, and intra-team types. To overcome the challenges of direct human feedback, we employ a Large Language Model (LLM) evaluator to validate our approach using feedback scenarios such as region constraints, collision avoidance, and task allocation. Our method effectively refines USV swarm policies, addressing key challenges in multi-agent systems while maintaining fairness and performance consistency.

多智能体强化学习无人艇人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。