arXiv:2602.07132stat.MLcs.LG2026-02被引 9

提出离散版伴随匹配,用于优化离散生成模型的奖励

Discrete Adjoint Matching

  • 从统计角度构建离散伴随估计器,解决非可导问题
  • 在数学推理任务中显著提升生成模型性能
  • 适合需要精细微调的离散生成模型研究者

求解熵正则化奖励优化问题的计算方法发展迅速,其中伴随匹配(AM)在连续状态空间和可导奖励下表现优异。然而,将其推广到离散生成建模仍面临巨大挑战,主要源于生成模型从连续到离散状态空间的转变,导致不可导性。本文提出离散伴随匹配(DAM),一种适用于基于连续时间马尔可夫链的离散生成模型(如扩散型大语言模型)的AM变体。DAM的核心是引入离散伴随——一个在离散域上定义的最优解估计器,使标准匹配框架得以应用。该方法基于纯统计视角,不同于AM的控制理论视角,为通用伴随估计器开辟新路径。我们在合成数据和数学推理任务中验证了DAM的有效性。

原文摘要 · Abstract (English)

Computation methods for solving entropy-regularized reward optimization -- a class of problems widely used for fine-tuning generative models -- have advanced rapidly. Among those, Adjoint Matching (AM, Domingo-Enrich et al., 2025) has proven highly effective in continuous state spaces with differentiable rewards. Transferring these practical successes to discrete generative modeling, however, remains particularly challenging and largely unexplored, mainly due to the drastic shift in generative model classes to discrete state spaces, which are nowhere differentiable. In this work, we propose Discrete Adjoint Matching (DAM) -- a discrete variant of AM for fine-tuning discrete generative models characterized by Continuous-Time Markov Chains, such as diffusion-based large language models. The core of DAM is the introduction of discrete adjoint-an estimator of the optimal solution to the original problem but formulated on discrete domains-from which standard matching frameworks can be applied. This is derived via a purely statistical standpoint, in contrast to the control-theoretic viewpoint in AM, thereby opening up new algorithmic opportunities for general adjoint-based estimators. We showcase DAM's effectiveness on synthetic and mathematical reasoning tasks.

生成模型离散优化伴随匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。