TCSM让离散扩散模型更灵活,可直接训练或用奖励微调。
Target Concrete Score Matching: A Holistic Framework for Discrete Diffusion
- 通过估计目标分布的原始数据空间得分,统一训练与微调流程。
- 在语言建模任务中性能优于或相当现有方法,样本效率更高。
- 适合需要预训练、奖励微调或知识蒸馏的离散数据生成场景。
离散扩散是建模和生成离散数据的有前景框架。本文提出目标具体得分匹配(TCSM),一种新颖且通用的离散扩散模型训练与微调目标。TCSM提供广泛适用的通用框架,支持直接从数据样本进行预训练,许多现有离散扩散方法自然成为其特例。此外,同一TCSM目标可扩展至后训练阶段,包括基于奖励函数或偏好数据的微调,以及从预训练自回归模型中进行知识蒸馏。这些新能力源于核心思想:估计目标分布在原始(干净)数据空间中的具体得分。这使得与仅在干净数据空间中操作的奖励函数和预训练模型无缝集成,而非依赖扩散过程的噪声中间空间。实验表明,TCSM在语言建模任务中表现匹配或超越当前方法,且兼具预训练与后训练适用性,灵活性更强,样本效率更高。
原文摘要 · Abstract (English)
Discrete diffusion is a promising framework for modeling and generating discrete data. In this work, we present Target Concrete Score Matching (TCSM), a novel and versatile objective for training and fine-tuning discrete diffusion models. TCSM provides a general framework with broad applicability. It supports pre-training discrete diffusion models directly from data samples, and many existing discrete diffusion approaches naturally emerge as special cases of our more general TCSM framework. Furthermore, the same TCSM objective extends to post-training of discrete diffusion models, including fine-tuning using reward functions or preference data, and distillation of knowledge from pre-trained autoregressive models. These new capabilities stem from the core idea of TCSM, estimating the concrete score of the target distribution, which resides in the original (clean) data space. This allows seamless integration with reward functions and pre-trained models, which inherently only operate in the clean data space rather than the noisy intermediate spaces of diffusion processes. Our experiments on language modeling tasks demonstrate that TCSM matches or surpasses current methods. Additionally, TCSM is versatile, applicable to both pre-training and post-training scenarios, offering greater flexibility and sample efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。