用新方法让语言生成快64倍,还更准。
Ultra-Fast Language Generation via Discrete Diffusion Divergence Instruct
- 基于积分KL散度最小化,从扩散模型蒸馏出快速生成的轻量模型。
- 在OpenWebText上达到18.4困惑度,比GPT-2和以往加速模型更好。
- 训练时间减少20倍以上,适合需要快速生成的部署场景。
高效且高质量的语言生成是当前AI领域的核心追求。本文提出基于训练的离散扩散发散指令(DiDi-Instruct)方法,从预训练的扩散大语言模型(dLLM)出发,蒸馏出少数步骤的轻量学生模型以实现快速生成。该方法在保持或超越教师模型与GPT-2基线性能的同时,实现最高达64倍的加速。其理论基础为基于积分KL散度最小化的新型框架,导出实用训练算法。通过引入分组奖励归一化、中间状态匹配及奖励引导的祖先采样器,显著提升训练稳定性、模型覆盖范围与推理质量。在OpenWebText基准上,DiDi-Instruct困惑度从8次非显式函数评估(NFE)的62.2降至128 NFE的18.4,优于已有加速dLLM与GPT-2基线。该方法仅带来约1%的熵损失,且额外训练耗时比竞争方法减少20倍以上。通过大量消融实验、模型缩放、下游任务评估及无条件蛋白质序列生成验证了其鲁棒性与有效性。结论:DiDi-Instruct实现了语言生成中高效且可靠的蒸馏。
原文摘要 · Abstract (English)
Fast and high-quality language generation is the holy grail that people pursue in the age of AI. In this work, we introduce Discrete Diffusion Divergence Instruct (DiDi-Instruct), a training-based method that initializes from a pre-trained diffusion large language model (dLLM) and distills a few-step student for fast generation. The model distilled with DiDi-Instruct matches or surpasses its dLLM teacher and the GPT-2 baseline while providing up to 64$\times$ acceleration. The theoretical foundation of DiDi-Instruct is a novel framework based on integral KL-divergence minimization, which leads to a practical training algorithm. We further introduce grouped reward normalization, intermediate-state matching, and the reward-guided ancestral sampler to improve training stability, model coverage, and inference quality. On the OpenWebText benchmark, DiDi-Instruct achieves perplexity ranging from 62.2 (8 NFEs) to 18.4 (128 NFEs), outperforming prior accelerated dLLMs and the GPT-2 baseline. These gains incur a negligible entropy loss (around $1$%) and reduce additional training wall-clock time by more than $20\times$ compared to competing dLLM distillation methods. We further validate the robustness and effectiveness of DiDi-Instruct through extensive ablation studies, model scaling, downstream task evaluations, and unconditional protein sequence generation. In conclusion, DiDi-Instruct enables efficient and effective distillation for language generation in the blink of an eye.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。