让大模型学会预判对手行为,提升多智能体决策能力
Foresight Optimization for Strategic Reasoning in Large Language Models

- 将对手建模融入策略优化,显式增强前瞻推理能力
- 在合作与竞争任务中显著提升模型战略表现
- 适用于需要预判与博弈的复杂场景,如游戏或谈判
大型语言模型的推理能力虽已大幅提升,但在多智能体环境中仍难以有效决策,主要因缺乏显式的前瞻性建模。为此,本文提出前瞻性策略优化(FoPO),将对手建模引入策略优化框架,使模型能同时考虑自身利益与对方影响,从而实现更优的战略推理。我们构建了两个精心设计的数据集:Cooperative RSA 和 Competitive Taboo,规则清晰、难度适中,支持在自对弈框架下系统评估 FoPO。实验表明,无论模型大小或来源,采用 FoPO 训练的模型均显著提升战略推理能力,且在跨领域战略场景中表现出强泛化性能,显著优于标准推理优化基线。
原文摘要 · Abstract (English)
Reasoning capabilities in large language models (LLMs) have generally advanced significantly. However, it is still challenging for existing reasoning-based LLMs to perform effective decision-making abilities in multi-agent environments, due to the absence of explicit foresight modeling. To this end, strategic reasoning, the most fundamental capability to anticipate the counterpart's behaviors and foresee its possible future actions, has been introduced to alleviate the above issues. Strategic reasoning is fundamental to effective decision-making in multi-agent environments, yet existing reasoning enhancement methods for LLMs do not explicitly capture its foresight nature. In this work, we introduce Foresight Policy Optimization (FoPO) to enhance strategic reasoning in LLMs, which integrates opponent modeling principles into policy optimization, thereby enabling explicit consideration of both self-interest and counterpart influence. Specifically, we construct two curated datasets, namely Cooperative RSA and Competitive Taboo, equipped with well-designed rules and moderate difficulty to facilitate a systematic investigation of FoPO in a self-play framework. Our experiments demonstrate that FoPO significantly enhances strategic reasoning across LLMs of varying sizes and origins. Moreover, models trained with FoPO exhibit strong generalization to out-of-domain strategic scenarios, substantially outperforming standard LLM reasoning optimization baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。