让大模型在扑克中学会长远布局,提升长期胜率。
Implicit Strategic Optimization: Rethinking Long-Horizon Decision-Making in Adversarial Poker Environments
- 用策略预测修正政策,实时调整决策思路。
- 6人德州扑克实验中,长期收益超越主流基线。
- 即使预测有误差,表现仍稳定,适合复杂博弈场景。
在对抗性扑克环境中,训练大型语言模型代理通常依赖短期目标(如胜率)。但在长周期设定下,收益受随时间演变的隐含策略外部性影响,导致短视优化和基于方差的后悔分析失效,即使动态可预测。为此,我们提出隐式策略优化(ISO),一种预测感知框架:每个代理预测当前策略上下文,并据此在线更新策略。ISO结合策略奖励模型(SRM)以估计动作的长期战略价值,以及iso-grpo——一种上下文相关的乐观学习规则。我们证明了亚线性上下文后悔和均衡收敛保证,主导项与上下文误判次数相关;当预测误差受限时,边界恢复已知策略外部性的静态博弈速率。在6人无限制德州扑克和竞争性宝可梦实验中,相比强基线模型,长期回报持续提升,且在可控预测噪声下表现平稳退化。
原文摘要 · Abstract (English)
Training large language model (LLM) agents for adversarial games is often driven by episodic objectives such as win rate. In long-horizon settings, however, payoffs are shaped by latent strategic externalities that evolve over time, so myopic optimization and variation-based regret analyses can become vacuous even when the dynamics are predictable. To solve this problem, we introduce Implicit Strategic Optimization (ISO), a prediction-aware framework in which each agent forecasts the current strategic context and uses it to update its policy online. ISO combines a Strategic Reward Model (SRM) that estimates the long-run strategic value of actions with iso-grpo, a context-conditioned optimistic learning rule. We prove sublinear contextual regret and equilibrium convergence guarantees whose dominant terms scale with the number of context mispredictions; when prediction errors are bounded, our bounds recover the static-game rates obtained when strategic externalities are known. Experiments in 6-player No-Limit Texas Hold'em and competitive Pokemon show consistent improvements in long-term return over strong LLM and RL baselines, and graceful degradation under controlled prediction noise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。