arXiv:2412.02016cs.LGcs.AI2024-12

将强化学习与博弈论结合,提升复杂环境下的均衡逼近能力。

Explore Reinforced: Equilibrium Approximation with Reinforcement Learning

  • 分离行动选择与均衡计算,保持学习完整性
  • 在网络安全和多臂老虎机中性能显著提升
  • 适合需要稳定策略的复杂对抗场景

当前近似粗相关均衡(CCE)算法在大型随机环境中难以有效逼近均衡,但理论上可收敛到强解。相比之下,现代强化学习(RL)算法训练更快,但解的质量较弱。我们提出Exp3-IXrl——融合RL与博弈论的方法,将RL代理的动作选择与均衡计算分离,同时保持学习过程的完整性。实验表明,该算法拓展了均衡逼近算法的应用范围,在复杂的对抗性网络安全环境(Cyber Operations Research Gym)及经典多臂老虎机设置中均表现更优。

原文摘要 · Abstract (English)

Current approximate Coarse Correlated Equilibria (CCE) algorithms struggle with equilibrium approximation for games in large stochastic environments but are theoretically guaranteed to converge to a strong solution concept. In contrast, modern Reinforcement Learning (RL) algorithms provide faster training yet yield weaker solutions. We introduce Exp3-IXrl - a blend of RL and game-theoretic approach, separating the RL agent's action selection from the equilibrium computation while preserving the integrity of the learning process. We demonstrate that our algorithm expands the application of equilibrium approximation algorithms to new environments. Specifically, we show the improved performance in a complex and adversarial cybersecurity network environment - the Cyber Operations Research Gym - and in the classical multi-armed bandit settings.

强化学习博弈论网络安全均衡逼近

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。