arXiv:2607.18719cs.MAcs.AI2026-07

让部分智能体听指令,其余自动补位,提升团队协作灵活性。

Strategy-Following Multi-Agent Deep Reinforcement Learning Considering Control Strategies Provided to Other Agents

论文配图:Strategy-Following Multi-Agent Deep Reinforcement Learning Considering Control Strategies Provided to Other Agents
图 1 · 摘自论文原文
  • 仅对关键智能体给指令,其他自动补足任务
  • 新方法使团队协作性能优于传统方法
  • 适合需人机协同的动态场景应用

本研究提出一种多智能体深度强化学习方法,使智能体在学习后可由人类管理者通过简单指令控制,且未受指令的智能体能根据其他智能体行为隐式补全整体任务。当前多智能体深度学习应用虽具潜力,但为实现广泛社会应用,需支持人类以简便方式响应环境与社会变化。即使无外部变化,学习到的协作结构也常不符人类预期,因此应能调整协作模式以匹配人类意图。已有研究尝试用简单指令控制智能体行为,但假设所有智能体均接收指令,效率低且不适用于优化协作设计。理想情况是仅特定智能体接受关键动作指令,其余智能体自动完成剩余任务。所提方法扩展了先前多智能体强化学习中的可控性工作,使未受指令的智能体能自适应弥补被忽略的任务与区域。实验表明,采用该方法的智能体能切换至新的协作结构,并取得比传统方法更优的性能。

原文摘要 · Abstract (English)

This study proposes a learning method for multi-agent systems that allows agents to be controlled through human manager instructions after learning and enables uninstructed agents to implicitly complement the overall work based on the actions of other agents. Multi-agent applications using deep learning have shown potential; thus, to achieve extensive social applications, humans should be able to control learned agents using simple methods to respond to environmental and social changes. Even without such changes, learned coordination often does not match the expectations of human managers, making it preferable to control coordination structures to match human intentions. Some studies have aimed to control agent behavior using simple instructions. However, they assumed that instructions are provided to all agents, which is time-consuming and not evident when designing a better cooperation regime. Ideally, specific agents should receive key action instructions, while others should automatically complete the remaining tasks. The proposed method, which extends previous work on controllability in multi-agent deep reinforcement learning, enables uninstructed agents to adaptively complement overlooked tasks and areas. The experimental results show that agents using the proposed method can shift to another cooperative structure and achieve better performance than those using conventional methods.

多智能体强化学习人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。