arXiv:2509.01057physics.soc-phcs.AI2025-09被引 3

用Q学习动态调整网络连接,让智能体自发形成合作集群。

Q-Learning-Driven Adaptive Rewiring for Cooperative Control in Heterogeneous Networks

  • 基于时序差分学习的自适应重连机制,结合历史互动优化策略与关系
  • 在幂律网络中实现三类行为模式,高约束下合作水平显著提升
  • 适合研究多智能体协同、复杂系统自组织及强化学习应用者

多智能体系统中的合作涌现是统计物理中的基本问题,微观学习规则驱动宏观集体行为转变。本文提出一种基于Q-learning的自适应重连方法,结合时序差分学习与网络重构,使智能体根据交互历史优化策略与社会连接。通过邻居特异性Q-learning,智能体发展出复杂伙伴关系管理策略,促进合作集群形成,实现合作与背叛区域的空间分离。在反映真实异构连接特征的幂律网络上,评估不同重连约束下的涌现行为,发现参数空间内呈现多样合作模式而非尖锐热力学相变。系统分析揭示三种行为范式:宽松态(低约束)快速形成合作簇,中间态对困境强度敏感,保守态(高约束)通过战略积累逐步优化网络结构。模拟显示,适度约束产生抑制合作的过渡区,而完全自适应重连通过系统探索有利网络配置提升合作水平。定量分析表明,增加重连频率驱动大规模簇形成,且簇大小服从幂律分布。结果建立理解复杂自适应系统中智能驱动合作模式形成的全新范式,揭示机器学习作为多智能体网络自发组织的替代驱动力。

原文摘要 · Abstract (English)

Cooperation emergence in multi-agent systems represents a fundamental statistical physics problem where microscopic learning rules drive macroscopic collective behavior transitions. We propose a Q-learning-based variant of adaptive rewiring that builds on mechanisms studied in the literature. This method combines temporal difference learning with network restructuring so that agents can optimize strategies and social connections based on interaction histories. Through neighbor-specific Q-learning, agents develop sophisticated partnership management strategies that enable cooperator cluster formation, creating spatial separation between cooperative and defective regions. Using power-law networks that reflect real-world heterogeneous connectivity patterns, we evaluate emergent behaviors under varying rewiring constraint levels, revealing distinct cooperation patterns across parameter space rather than sharp thermodynamic transitions. Our systematic analysis identifies three behavioral regimes: a permissive regime (low constraints) enabling rapid cooperative cluster formation, an intermediate regime with sensitive dependence on dilemma strength, and a patient regime (high constraints) where strategic accumulation gradually optimizes network structure. Simulation results show that while moderate constraints create transition-like zones that suppress cooperation, fully adaptive rewiring enhances cooperation levels through systematic exploration of favorable network configurations. Quantitative analysis reveals that increased rewiring frequency drives large-scale cluster formation with power-law size distributions. Our results establish a new paradigm for understanding intelligence-driven cooperation pattern formation in complex adaptive systems, revealing how machine learning serves as an alternative driving force for spontaneous organization in multi-agent networks.

多智能体合作涌现Q学习网络重构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。