用扩散模型解决多智能体环境协同设计的可扩展性难题
Scaling Multi-Agent Environment Co-Design with Diffusion Models
- 引入投影通用引导技术,高效生成满足约束的最优环境
- 通过评论器蒸馏机制,实现智能体策略与环境的动态适配
- 在仓储等场景中减少66%仿真样本,奖励提升39%
智能体-环境协同设计通过联合优化智能体策略与环境配置以提升系统性能,在仓储物流、风电场管理等领域具有重要应用前景。然而现有方法难以扩展:在高维环境设计空间下失效,且在联合优化过程中面临样本效率低的问题。本文提出扩散协同设计(DiCoDe)框架,包含两项核心创新:一是提出投影通用引导(PUG)采样技术,可在满足空间分离等硬约束条件下探索最大化奖励的环境分布;二是设计评论器蒸馏机制,将强化学习评论器的知识传递给扩散模型,使其能根据演化中的智能体策略获取密集且实时的学习信号。在仓储自动化、多智能体路径规划和风电场优化等复杂基准测试中,该方法持续超越现有最先进水平,例如在仓储场景中奖励提升39%,仿真样本减少66%。这为协同设计在真实世界中的落地树立了新标准。
原文摘要 · Abstract (English)
The agent-environment co-design paradigm jointly optimises agent policies and environment configurations in search of improved system performance. With application domains ranging from warehouse logistics to windfarm management, co-design promises to fundamentally change how we deploy multi-agent systems. However, current co-design methods struggle to scale. They collapse under high-dimensional environment design spaces and suffer from sample inefficiency when addressing moving targets inherent to joint optimisation. We address these challenges by developing Diffusion Co-Design (DiCoDe), a scalable and sample-efficient co-design framework pushing co-design towards practically relevant settings. DiCoDe incorporates two core innovations. First, we introduce Projected Universal Guidance (PUG), a sampling technique that enables DiCoDe to explore a distribution of reward-maximising environments while satisfying hard constraints such as spatial separation between obstacles. Second, we devise a critic distillation mechanism to share knowledge from the reinforcement learning critic, ensuring that the guided diffusion model adapts to evolving agent policies using a dense and up-to-date learning signal. Together, these improvements lead to superior environment-policy pairs when validated on challenging multi-agent environment co-design benchmarks including warehouse automation, multi-agent pathfinding and wind farm optimisation. Our method consistently exceeds the state-of-the-art, achieving, for example, 39% higher rewards in the warehouse setting with 66% fewer simulation samples. This sets a new standard in agent-environment co-design, and is a stepping stone towards reaping the rewards of co-design in real world domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。