arXiv:2604.27166cs.LGcs.GT2026-04被引 1

通过博弈论框架实现答案级微调,提升语言模型的准确性与一致性。

Distributional Alignment Games for Answer-Level Fine-Tuning

  • 构建生成器与目标分布的博弈模型,将复杂优化转为可解投影问题。
  • 在数学推理任务中,相比传统方法显著降低计算复杂度。
  • 统一了多样性与自提升机制,适合需要高准确性的问答系统。

我们聚焦于答案级微调(ALFT)问题,目标是根据语言模型最终答案的正确性或属性进行优化,而非依赖具体的推理路径。直接优化答案级目标因需对庞大的潜在推理路径空间进行边际化而计算上不可行。为此,我们提出一种通用的博弈论框架,将问题提升至分布对齐博弈。我们将ALFT建模为生成器(策略)与辅助分布(目标)之间的双人博弈,并证明该博弈的纳什均衡恰好对应原始答案级优化问题的解。这一变分视角将难以处理的边际化问题转化为可解的投影问题。我们证明该框架统一了近期关于多样性和自我改进(连贯性)的方法,并提供了与组相对策略优化(GRPO)兼容的高效算法,如连贯性-GRPO,显著降低了数学推理任务中的计算复杂度。

原文摘要 · Abstract (English)

We focus on the problem of \emph{Answer-Level Fine-Tuning} (ALFT), where the goal is to optimize a language model based on the correctness or properties of its final answers, rather than the specific reasoning traces used to produce them. Directly optimizing answer-level objectives is computationally intractable due to the need to marginalize over the vast space of latent reasoning paths. To overcome this, we propose a general game-theoretical framework that lifts the problem to a \emph{Distributional Alignment Game}. We formulate ALFT as a two-player game between a Policy (the generator) and a Target (an auxiliary distribution). We prove that the Nash Equilibrium of this game corresponds exactly to the solution of the original answer-level optimization problem. This variational perspective transforms the intractable marginalization problem into a tractable projection problem. We demonstrate that this framework unifies recent approaches to diversity and self-improvement (coherence) and provide efficient algorithms compatible with Group Relative Policy Optimization (GRPO), such as Coherence-GRPO, yielding significant complexity gains in mathematical reasoning tasks.

语言模型微调博弈论推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。