用偏好数据直接优化离散扩散模型,无需奖励函数。
Preference-Based Alignment of Discrete Diffusion Models
- 将DPO思想适配到离散扩散模型,基于偏好数据直接调整生成过程。
- 在二进制序列生成任务中实现偏好对齐,保持结构有效性。
- 无需显式奖励模型,适合语言与蛋白序列等复杂生成任务。
扩散模型在多个领域达到顶尖性能,近期进展已将其拓展至离散数据。然而,在缺乏显式奖励函数的情况下,如何将离散扩散模型与特定任务偏好对齐仍具挑战。本文提出首个将直接偏好优化(DPO)适配到离散扩散模型的方法——离散扩散DPO(D2-DPO),该模型被形式化为连续时间马尔可夫链。我们推导出一种新型损失函数,利用偏好数据直接微调生成过程,同时保持与参考分布的一致性。我们在结构化二进制序列生成任务上验证了D2-DPO的有效性,结果表明该方法能有效对齐模型输出与偏好,同时维持结构合理性。实验表明,D2-DPO可在不依赖显式奖励模型的前提下实现可控微调,是强化学习方法的实用替代方案。未来工作将探索其在语言建模和蛋白质序列生成中的扩展,并研究如均匀加噪等替代噪声策略以提升应用灵活性。
原文摘要 · Abstract (English)
Diffusion models have achieved state-of-the-art performance across multiple domains, with recent advancements extending their applicability to discrete data. However, aligning discrete diffusion models with task-specific preferences remains challenging, particularly in scenarios where explicit reward functions are unavailable. In this work, we introduce Discrete Diffusion DPO (D2-DPO), the first adaptation of Direct Preference Optimization (DPO) to discrete diffusion models formulated as continuous-time Markov chains. Our approach derives a novel loss function that directly fine-tunes the generative process using preference data while preserving fidelity to a reference distribution. We validate D2-DPO on a structured binary sequence generation task, demonstrating that the method effectively aligns model outputs with preferences while maintaining structural validity. Our results highlight that D2-DPO enables controlled fine-tuning without requiring explicit reward models, making it a practical alternative to reinforcement learning-based approaches. Future research will explore extending D2-DPO to more complex generative tasks, including language modeling and protein sequence generation, as well as investigating alternative noise schedules, such as uniform noising, to enhance flexibility across different applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。