arXiv:2604.09028cs.MAcs.LG2026-04被引 2

无人机应急通信中,新算法让多智能体策略自适应变化,抗干扰更强。

Plasticity-Enhanced Multi-Agent Mixture of Experts for Dynamic Objective Adaptation in UAVs-Assisted Emergency Communication Networks

论文配图:Plasticity-Enhanced Multi-Agent Mixture of Experts for Dynamic Objective Adaptation in UAVs-Assisted Emergency Communication Networks
图 1 · 摘自论文原文
  • 每个无人机用专家混合模型,动态选专用策略执行
  • 切换阶段时注入扰动,提升策略灵活性,性能提升26.3%
  • 适合高动态用户场景的应急通信系统部署

无人飞行器作为空中基站可在灾后快速恢复通信,但用户移动性和流量需求的突变导致服务质量权衡剧烈变化,引发强非平稳性。深度强化学习策略在该环境下易出现塑性损失,表现为表征坍塌和神经元休眠,影响适应能力。本文提出塑性增强的多智能体专家混合模型(PE-MAMoE),采用集中训练、分散执行框架,基于多智能体近端策略优化。每个无人机配备稀疏门控的专家混合动作器,路由器每步仅激活一个专家。引入无参相位控制器,在相位切换后施加短暂的专家专属随机扰动,重置动作对数标准差,降低熵与学习率,并调度路由器温度,以恢复策略可塑性而不破坏安全行为。理论推导出动态后悔上界,表明跟踪误差与环境变化及累积噪声能量相关。在含移动用户与3GPP式信道的相位驱动仿真中,PE-MAMoE相较最优基线提升归一化四分位均值回报26.3%,服务用户容量增加12.8%,碰撞减少约75%。诊断分析证实,在阶段切换时专家特征秩持续更高,且周期性唤醒休眠神经元。

原文摘要 · Abstract (English)

Unmanned aerial vehicles serving as aerial base stations can rapidly restore connectivity after disasters, yet abrupt changes in user mobility and traffic demands shift the quality of service trade-offs and induce strong non-stationarity. Deep reinforcement learning policies suffer from plasticity loss under such shifts, as representation collapse and neuron dormancy impair adaptation. We propose plasticity enhanced multi-agent mixture of experts (PE-MAMoE), a centralized training with decentralized execution framework built on multi-agent proximal policy optimization. PE-MAMoE equips each UAV with a sparsely gated mixture of experts actor whose router selects a single specialist per step. A non-parametric Phase Controller injects brief, expert-only stochastic perturbations after phase switches, resets the action log-standard-deviation, anneals entropy and learning rate, and schedules the router temperature, all to re-plasticize the policy without destabilizing safe behaviors. We derive a dynamic regret bound showing the tracking error scales with both environment variation and cumulative noise energy. In a phase-driven simulator with mobile users and 3GPP-style channels, PE-MAMoE improves normalized interquartile mean return by 26.3\% over the best baseline, increases served-user capacity by 12.8\%, and reduces collisions by approximately 75\%. Diagnostics confirm persistently higher expert feature rank and periodic dormant-neuron recovery at regime switches.

无人机通信强化学习动态适应多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。