提出A2D2框架,实现任意长度文本生成的自适应解码优化。
A2D2: Fine-Tuning Any-Length Discrete Diffusion for Adaptive Decoding

- 联合优化插入与去掩码策略,实现奖励引导的统一微调。
- 理论证明可收敛至不可计算的奖励倾斜分布,无需目标样本。
- 支持灵活长度生成,提升准确率与解码效率,适合长序列任务。
离散扩散模型为序列生成提供了简单稳定的似然框架,最近通过词元插入扩展至任意长度设置。然而,针对任意长度离散扩散模型的原理性奖励引导微调仍缺乏探索。本文提出自适应解码的任意长度离散扩散微调框架(A2D2),通过联合优化插入与去掩码策略,并结合基于质量的推理调度,实现奖励引导的统一微调。我们推导了联合插入-去掩码路径测度的Radon-Nikodym导数,理论上保证在无需目标样本的情况下收敛至不可计算的奖励倾斜序列分布。在此基础上,我们将去掩码与插入质量作为可计算的解码误差最小化方法,提出自适应联合解码(AJD)损失,其可证明地生成最优路径测度以逼近奖励倾斜分布。实验表明,A2D2在奖励优化、生成灵活性和准确性方面均优于先前的固定长度微调与推理时引导方法。
原文摘要 · Abstract (English)
Discrete diffusion models offer a simple and stable likelihood-based framework for sequence generation, recently extended to any-length settings via token insertion. Principled reward-guided fine-tuning for any-length discrete diffusion, however, remains largely unexplored. We introduce Fine-Tuning Any-Length Discrete Diffusion for Adaptive Decoding (A2D2), a unified framework for reward-guided fine-tuning of any-length discrete diffusion models via joint optimization of the insertion and unmasking policies together with a quality-based inference schedule. We derive the Radon-Nikodym derivative for the joint insertion-unmasking path measures, enabling theoretically guaranteed convergence to the intractable reward-tilted sequence distribution without requiring target samples. Building on this, we establish unmasking and insertion quality as tractable approaches for minimizing decoding error and introduce the Adaptive Joint Decoding (AJD) loss, which provably yields the optimal path measure that generates the reward-tilted distribution. Empirically, A2D2 improves reward optimization while enhancing generation flexibility and accuracy over prior fixed-length fine-tuning and inference-time guidance methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。