arXiv:2510.23658cs.LGcs.AI2025-10被引 2

无需配对数据,用新方法让扩散语言模型更符合人类偏好。

Aligning Diffusion Language Models via Unpaired Preference Optimization

  • 用变分下界替代难以计算的扩散模型似然,结合前景理论优化偏好。
  • 在两个评测集上分别达到65.9%和62.3%的调整胜率,优于基线模型。
  • 适合研究扩散模型对齐、偏好学习及减少人工标注成本的团队。

扩散语言模型(dLLMs)是自回归生成器的新兴替代方案,但将其对齐至人类偏好面临挑战:序列似然难以计算,且成对偏好数据收集成本高。本文提出ELBO-KTO,将期望下界(ELBO)作为扩散模型似然的代理,并结合前景理论驱动的无配对偏好目标(Kahneman-Tversky优化,KTO)。我们分析了ELBO替换带来的偏差与方差,并采用方差缩减策略稳定训练梯度。应用于LLaDA-8B-Instruct时,ELBO-KTO在kto-mix-14k和UltraFeedback-Binary评测集上分别获得65.9%和62.3%的调整胜率,优于基线模型(使用自动大模型评判)。在下游任务如GSM8K、MMLU及其他推理/知识基准测试中,该方法在相同解码条件下表现持平或超越基线模型。这证明无配对偏好优化是扩散语言模型对齐的可行替代方案。

原文摘要 · Abstract (English)

Diffusion language models (dLLMs) are an emerging alternative to autoregressive (AR) generators, but aligning them to human preferences is challenging because sequence log-likelihoods are intractable and pairwise preference data are costly to collect. We introduce ELBO-KTO, which combines an ELBO surrogate for diffusion log-likelihoods with a prospect-theoretic, unpaired preference objective (Kahneman Tversky Optimization, KTO). We analyze the bias and variance induced by the ELBO substitution and employ variance-reduction practices that stabilize gradients during training. Applied to LLaDA-8B-Instruct, ELBO-KTO yields 65.9% and 62.3% adjusted win rates on kto-mix-14k and UltraFeedback-Binary, respectively, versus the base model under an automatic LLM judge. Across downstream tasks, including GSM8K, MMLU, and additional reasoning/knowledge benchmarks, ELBO-KTO trained on UltraFeedback-Binary performs on par with or better than the base model under identical decoding. This establishes unpaired preference optimization as a viable alternative to pairwise alignment in diffusion LLMs.

扩散模型偏好对齐无配对优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。