arXiv:2510.16761cs.CL2025-10

通过自对弈提升语言智能体在对抗游戏中的策略推理能力

Enhancing Language Agent Strategic Reasoning through Self-Play in Adversarial Games

  • 设计了基于自对弈的学习框架,无需专家标注数据
  • 自对弈使胜率相比基线平均提升30%,对GPT-4达54.76%
  • 适用于需要复杂策略博弈的AI系统研发

现有语言智能体在动态对抗游戏中常因策略推理不足而表现不佳。为解决此问题,我们提出一种无需依赖昂贵专家标注数据的自主学习方法——步骤级策略优化自学习框架(SCO-PAL)。该方法通过设置不同水平对手进行对抗实验,发现自对弈是提升策略推理最有效的方式。在六场对抗游戏中,采用自对弈的SCO-PAL相较基线平均胜率提升约30%,对GPT-4的胜率达54.76%。研究揭示了对手选择对学习性能的关键影响,推动了对抗环境中智能体训练机制的发展。

原文摘要 · Abstract (English)

Existing language agents often encounter difficulties in dynamic adversarial games due to poor strategic reasoning. To mitigate this limitation, a promising approach is to allow agents to learn from game interactions automatically, without relying on costly expert-labeled data. Unlike static environments where agents receive fixed feedback or rewards, selecting appropriate opponents in dynamic adversarial games can significantly impact learning performance. However, the discussion of opponents in adversarial environments remains an area under exploration. In this paper, we propose a Step-level poliCy Optimization method through Play-And-Learn, SCO-PAL. Leveraging SCO-PAL, we conduct a detailed analysis of opponent selection by setting opponents at different levels and find that self-play is the most effective way to improve strategic reasoning in such adversarial environments. Utilizing SCO-PAL with self-play, we increase the average win rate against four opponents by approximately 30% compared to baselines and achieve a 54.76% win rate against GPT-4 in six adversarial games.

语言智能体自对弈策略推理对抗游戏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。