多智能体牧羊任务中,用强化学习让机器人自动驱赶分散目标
Hierarchical Policy-Gradient Reinforcement Learning for Multi-Agent Shepherding Control of Non-Cohesive Targets
- 用近端策略优化统一选择和驱赶目标,避免动作离散化问题
- 无需先验动力学知识,可处理更多目标且感知受限时仍有效
- 适合需要自主协作的多机器人系统,如搜救或运输场景
我们提出一种去中心化的强化学习方法,用于多智能体对非聚集目标的牧羊控制,采用近端策略优化(Proximal Policy Optimization)实现目标选择与目标驱动的融合,克服了以往深度Q网络方法在离散动作上的限制,使智能体轨迹更平滑。该无模型框架在无需先验动力学知识的情况下有效解决牧羊问题。实验表明,该方法在目标数量增加及感知能力受限条件下仍具备有效性与可扩展性。
原文摘要 · Abstract (English)
We propose a decentralized reinforcement learning solution for multi-agent shepherding of non-cohesive targets using policy-gradient methods. Our architecture integrates target-selection with target-driving through Proximal Policy Optimization, overcoming discrete-action constraints of previous Deep Q-Network approaches and enabling smoother agent trajectories. This model-free framework effectively solves the shepherding problem without prior dynamics knowledge. Experiments demonstrate our method's effectiveness and scalability with increased target numbers and limited sensing capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。