arXiv:2506.23734cs.NEcs.AI2025-06

通过基因标记机制稳定对抗进化,提升复杂博弈中的协作稳定性。

Governing Strategic Dynamics: Equilibrium Stabilization via Divergence-Driven Control

  • 用跨代标记个体锚定评估,减少噪声干扰导致的策略震荡
  • 在猎鹿博弈和资源枯竭马尔可夫游戏中实现接近1的协作概率
  • 适合研究多智能体协同、对抗训练不稳定的场景

黑箱共演化在混合动机博弈中常受对手漂移非平稳性和噪声回放影响,扭曲进展信号并引发循环、红皇后效应与脱离。本文提出标记基因法(MGM),一种受课程启发的治理机制,通过跨代标记个体锚定评估,并结合DWAM与保守更新规则减少虚假更新。还引入NGD-Div,利用发散代理与自然梯度优化动态调整关键更新阈值。理论分析基于严格竞争环境,在协调博弈与资源耗尽马尔可夫游戏中评估了集成进化策略的MGM-E-NES。该方法在猎鹿博弈和性别之战中可靠恢复目标协作,最终协作概率接近(1,1)(如0.991±0.01/1.00±0.00 和 0.97±0.00/0.97±0.00)。在马尔可夫资源游戏中,30次种子实验下保持高且稳定的条件协作,最终协作率约为0.954/0.980/0.916(双方;标准差小),体现福利对齐与状态依赖行为。整体上,MGM-E-NES在任务间迁移仅需微调超参数,训练动态一致稳定,表明顶层治理可显著提升黑箱共演化在动态环境中的鲁棒性。

原文摘要 · Abstract (English)

Black-box coevolution in mixed-motive games is often undermined by opponent-drift non-stationarity and noisy rollouts, which distort progress signals and can induce cycling, Red-Queen dynamics, and detachment. We propose the \emph{Marker Gene Method} (MGM), a curriculum-inspired governance mechanism that stabilizes selection by anchoring evaluation to cross-generational marker individuals, together with DWAM and conservative marker-update rules to reduce spurious updates. We also introduce NGD-Div, which adapts the key update threshold using a divergence proxy and natural-gradient optimization. We provide theoretical analysis in strictly competitive settings and evaluate MGM integrated with evolution strategies (MGM-E-NES) on coordination games and a resource-depletion Markov game. MGM-E-NES reliably recovers target coordination in Stag Hunt and Battle of the Sexes, achieving final cooperation probabilities close to $(1,1)$ (e.g., $0.991\pm0.01/1.00\pm0.00$ and $0.97\pm0.00/0.97\pm0.00$ for the two players). In the Markov resource game, it maintains high and stable state-conditioned cooperation across 30 seeds, with final cooperation of $\approx 0.954/0.980/0.916$ in \textsc{Rich}/\textsc{Poor}/\textsc{Collapsed} (both players; small standard deviations), indicating welfare-aligned and state-dependent behavior. Overall, MGM-E-NES transfers across tasks with minimal hyperparameter changes and yields consistently stable training dynamics, showing that top-level governance can substantially improve the robustness of black-box coevolution in dynamic environments.

多智能体博弈论协同进化稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。