让大模型通过游戏自博弈学会可迁移的推理能力
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play
- 用轨迹调制机制筛选抽象推理路径,避免陷入具体游戏规则
- 在数学推理等任务上显著提升,竞赛级数学题效果尤佳
- 适合研究通用推理、强化学习与模型可迁移性的学者
游戏为语言模型发展通用推理能力提供了有力范式,因其天然要求战略规划、概率推断与适应性决策。然而,现有自博弈方法仅依赖最终游戏结果,无法区分可迁移的推理模式与仅适用于特定游戏的启发式策略。我们提出STRATAGEM,克服两个核心障碍:领域特异性(学习模式被锚定在游戏语义中)与上下文静止性(静态游戏环境难以促进推理演化)。STRATAGEM通过推理可迁移系数选择性强化展现抽象、跨域推理的轨迹,并通过推理演化奖励激励推理能力的逐步发展。在数学推理、通用推理及代码生成等多个基准测试中均取得显著提升,尤其在需多步推理的竞赛级数学任务中表现突出。消融实验与人工评估证实,两项机制均对可迁移推理有贡献。
原文摘要 · Abstract (English)
Games offer a compelling paradigm for developing general reasoning capabilities in language models, as they naturally demand strategic planning, probabilistic inference, and adaptive decision-making. However, existing self-play approaches rely solely on terminal game outcomes, providing no mechanism to distinguish transferable reasoning patterns from game-specific heuristics. We present STRATAGEM, which addresses two fundamental barriers to reasoning transfer: domain specificity, where learned patterns remain anchored in game semantics, and contextual stasis, where static game contexts fail to cultivate progressive reasoning. STRATAGEM selectively reinforces trajectories exhibiting abstract, domain-agnostic reasoning through a Reasoning Transferability Coefficient, while incentivizing adaptive reasoning development via a Reasoning Evolution Reward. Experiments across mathematical reasoning, general reasoning, and code generation benchmarks demonstrate substantial improvements, with particularly strong gains on competition-level mathematics where multi-step reasoning is critical. Ablation studies and human evaluation confirm that both components contribute to transferable reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。