提出新采样方法提升生成效率与压缩性能
List-Level Distribution Coupling with Applications to Speculative Decoding and Lossy Compression
- 基于广义Gumbel-max采样构建列表级耦合机制
- 在语言生成任务中实现媲美基线的解码速度与稳定性
- 适用于需要高效采样与低损压缩的场景
我们研究概率分布耦合问题的一种松弛形式:从一个分布生成一组样本,若其中任意一个样本与另一个分布生成的样本相同,则判定为接受。本文提出一种新型采样方法,扩展了Daliri等人(arXiv:2408.07978)提出的Gumbel-max采样以实现分布耦合,并建立了相应的接受概率下界,称为列表匹配引理。随后,我们讨论了两个应用:第一,设计了一种新的多草稿推测采样机制,实现简单且在多种语言任务中表现优于SpecTr和SpecInfer等基线方法;该方法还保证输出标记的某种草稿不变性,这是现有方案不支持的特性,并给出了令牌级接受概率的理论下界。第二,在存在侧信息的分布式有损压缩场景中,源样本被压缩后供多个解码器使用,每个解码器具有独立的侧信息。我们基于广义Gumbel-max采样提出一种压缩技术,在合成高斯源和MNIST图像数据集上的实验中均显示出显著增益。
原文摘要 · Abstract (English)
We study a relaxation of the problem of coupling probability distributions -- a list of samples is generated from one distribution and an accept is declared if any one of these samples is identical to the sample generated from the other distribution. We propose a novel method for generating samples, which extends the Gumbel-max sampling suggested in Daliri et al. (arXiv:2408.07978) for coupling probability distributions. We also establish a corresponding lower bound on the acceptance probability, which we call the list matching lemma. We next discuss two applications of our setup. First, we develop a new mechanism for multi-draft speculative sampling that is simple to implement and achieves performance competitive with baselines such as SpecTr and SpecInfer across a range of language tasks. Our method also guarantees a certain degree of drafter invariance with respect to the output tokens which is not supported by existing schemes. We also provide a theoretical lower bound on the token level acceptance probability. As our second application, we consider distributed lossy compression with side information in a setting where a source sample is compressed and available to multiple decoders, each with independent side information. We propose a compression technique that is based on our generalization of Gumbel-max sampling and show that it provides significant gains in experiments involving synthetic Gaussian sources and the MNIST image dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。