解决无人机应急通信中多智能体强化学习的神经元失效问题
PRIME: Plasticity Recovery in Multi-Agent Environments for UAV-Assisted Emergency Communication Networks

- 通过双向检测激活与梯度状态,精准识别需重置的静默神经元
- 在动态环境中使平均回报提升24.9%,静默神经元比例降至10–20%
- 适用于需要长期稳定学习的协同多智能体系统,如应急通信网络
现有强化学习控制器多假设环境平稳,少数处理变化的模型仅响应外部环境而忽视网络内部状态。我们发现持续非平稳性会直接损害内部状态:目标变化导致神经元逐步失活,共享策略丧失学习能力。传统重置方法在共享参数多智能体训练中不安全,因部分看似静默的神经元仍接收强梯度信号,其活跃性依赖于所处理的智能体观测。PRIME(Plasticity Recovery In Multi-agent Environments)在干预前同时验证正向激活与反向梯度信号。将双向静默神经元框架扩展至合作多智能体强化学习,通过全队批量聚合激活与梯度统计,利用训练损失已沉积的反向信号而非人工构造代理信号,并仅重置同时满足激活静默与梯度沉默的神经元。有效表示得以保留,学习能力恢复。在阶段切换的无人机应急通信仿真器上,PRIME相比MAPPO提升四分位均值回报24.9%,静默神经元占比维持在10–20%(对比40–45%);消融实验表明收益主要来自梯度信号与团队级聚合,而非特定重置算子。动态遗憾界显示扰动代价随小静默子空间维度增长,而非完整参数量。
原文摘要 · Abstract (English)
Most reinforcement learning controllers for these networks assume stationary conditions, and the few that handle change react to the external environment while leaving the network's internal state unexamined. We show that sustained non-stationarity damages this internal state directly: as objectives shift, neurons progressively fall dormant and the shared policy loses the capacity to learn. The obvious remedy, resetting dormant neurons, is unsafe under shared-parameter multi-agent training: many neurons that appear inactive are still receiving strong training gradients, and whether a neuron appears dormant depends on which agent's observations it processes. PRIME (Plasticity Recovery In Multi-agent Environments) therefore verifies both directions before intervening. Extending the bidirectional Silent Neuron framework to cooperative multi-agent reinforcement learning, it aggregates activation and gradient statistics over the full team batch, reads the backward signal from the gradient the training loss has already deposited , not from a hand-crafted proxy, and reinitializes only neurons that are simultaneously activation-dormant and gradient-silent. Useful representations are preserved while learning capacity is restored. On a phase-switching UAV emergency communication simulator, PRIME improves interquartile mean return by 24.9\% over MAPPO and holds dormant neuron fractions at 10--20\% versus 40--45\%; ablations attribute the gains to the gradient signal and team-level aggregation rather than to the specific reset operator. A dynamic regret bound shows that the perturbation cost scales with the small silent-subspace dimension rather than the full parameter count.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。