arXiv:2602.08041cs.LGcs.AI2026-02被引 1

让大模型在扑克中学会长远布局,提升长期胜率。

Implicit Strategic Optimization: Rethinking Long-Horizon Decision-Making in Adversarial Poker Environments

  • 用策略预测修正政策,实时调整决策思路。
  • 6人德州扑克实验中,长期收益超越主流基线。
  • 即使预测有误差,表现仍稳定,适合复杂博弈场景。

在对抗性扑克环境中,训练大型语言模型代理通常依赖短期目标(如胜率)。但在长周期设定下,收益受随时间演变的隐含策略外部性影响,导致短视优化和基于方差的后悔分析失效,即使动态可预测。为此,我们提出隐式策略优化(ISO),一种预测感知框架:每个代理预测当前策略上下文,并据此在线更新策略。ISO结合策略奖励模型(SRM)以估计动作的长期战略价值,以及iso-grpo——一种上下文相关的乐观学习规则。我们证明了亚线性上下文后悔和均衡收敛保证,主导项与上下文误判次数相关;当预测误差受限时,边界恢复已知策略外部性的静态博弈速率。在6人无限制德州扑克和竞争性宝可梦实验中,相比强基线模型,长期回报持续提升,且在可控预测噪声下表现平稳退化。

原文摘要 · Abstract (English)

Training large language model (LLM) agents for adversarial games is often driven by episodic objectives such as win rate. In long-horizon settings, however, payoffs are shaped by latent strategic externalities that evolve over time, so myopic optimization and variation-based regret analyses can become vacuous even when the dynamics are predictable. To solve this problem, we introduce Implicit Strategic Optimization (ISO), a prediction-aware framework in which each agent forecasts the current strategic context and uses it to update its policy online. ISO combines a Strategic Reward Model (SRM) that estimates the long-run strategic value of actions with iso-grpo, a context-conditioned optimistic learning rule. We prove sublinear contextual regret and equilibrium convergence guarantees whose dominant terms scale with the number of context mispredictions; when prediction errors are bounded, our bounds recover the static-game rates obtained when strategic externalities are known. Experiments in 6-player No-Limit Texas Hold'em and competitive Pokemon show consistent improvements in long-term return over strong LLM and RL baselines, and graceful degradation under controlled prediction noise.

策略优化强化学习大模型博弈扑克游戏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。