arXiv:2512.03211cs.LGcs.NI2025-12被引 1

多智能体强化学习让路由器自主协作,提升网络传输效率。

A Multi-Agent, Policy-Gradient approach to Network Routing

  • 多个路由器作为智能体,通过策略梯度算法自主学习协作路由。
  • 优化奖励信号后,收敛速度显著提升,整体传输延迟降低。
  • 适合研究分布式智能决策、网络优化与强化学习应用的读者。

网络路由是典型的分布式决策问题,具有明确的性能指标,如数据包从源到目的地的平均传输时间。采用基于策略梯度的强化学习算法OLPOMDP,在多种网络模型的模拟中成功实现了路由优化。多个分布式智能体(路由器)在无显式通信情况下学会协同行为,避免了对个体有利但损害整体性能的策略。此外,通过显式惩罚特定次优行为模式来设计奖励信号,被证实可显著加快算法收敛速度。

原文摘要 · Abstract (English)

Network routing is a distributed decision problem which naturally admits numerical performance measures, such as the average time for a packet to travel from source to destination. OLPOMDP, a policy-gradient reinforcement learning algorithm, was successfully applied to simulated network routing under a number of network models. Multiple distributed agents (routers) learned co-operative behavior without explicit inter-agent communication, and they avoided behavior which was individually desirable, but detrimental to the group's overall performance. Furthermore, shaping the reward signal by explicitly penalizing certain patterns of sub-optimal behavior was found to dramatically improve the convergence rate.

多智能体强化学习网络路由

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。