解决城市系统中动态变数智能体的协同调度问题
Adaptive Value Decomposition: Coordinating a Varying Number of Agents in Urban Systems
- 动态适应智能体数量变化,支持异步决策
- 通过轻量机制缓解共享策略导致的行为同质化
- 在伦敦和华盛顿真实共享单车调度任务中表现更优
多智能体强化学习(MARL)为多智能体系统(MAS)协调提供了有前景的范式。然而,现有方法通常依赖固定智能体数量和完全同步动作执行等限制性假设,这在城市系统中常不成立——智能体数量随时间变化,且动作持续时间各异,形成半马尔可夫多智能体学习(semi-MARL)场景。此外,虽共享策略参数可提升学习效率,但在相似观测下部分智能体并发决策时,易导致行为高度同质化,降低协作质量。为此,本文提出自适应价值分解(AVD),一种能适应动态智能体群体的协作式MARL框架。AVD引入轻量级机制缓解共享策略引发的动作同质化,促进行为多样性并维持高效协作。同时设计了适配半MARL场景的训练-执行策略,支持异步决策。在伦敦和华盛顿特区的真实共享单车再分配任务中,实验表明AVD显著优于现有先进基线,验证了其有效性与泛化能力。
原文摘要 · Abstract (English)
Multi-agent reinforcement learning (MARL) provides a promising paradigm for coordinating multi-agent systems (MAS). However, most existing methods rely on restrictive assumptions, such as a fixed number of agents and fully synchronous action execution. These assumptions are often violated in urban systems, where the number of active agents varies over time, and actions may have heterogeneous durations, resulting in a semi-MARL setting. Moreover, while sharing policy parameters among agents is commonly adopted to improve learning efficiency, it can lead to highly homogeneous actions when a subset of agents make decisions concurrently under similar observations, potentially degrading coordination quality. To address these challenges, we propose Adaptive Value Decomposition (AVD), a cooperative MARL framework that adapts to a dynamically changing agent population. AVD further incorporates a lightweight mechanism to mitigate action homogenization induced by shared policies, thereby encouraging behavioral diversity and maintaining effective cooperation among agents. In addition, we design a training-execution strategy tailored to the semi-MARL setting that accommodates asynchronous decision-making when some agents act at different times. Experiments on real-world bike-sharing redistribution tasks in two major cities, London and Washington, D.C., demonstrate that AVD outperforms state-of-the-art baselines, confirming its effectiveness and generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。