arXiv:2503.05330cs.CLcs.AI2025-03EMNLP被引 5

提出新方法,让多路径推理更高效。

Speculative Decoding for Multi-Sample Inference

  • 利用多路径生成的共识,不依赖额外模型生成高质量草稿。
  • 数学推理任务上草稿接受率显著提升,构建草稿延迟更低。
  • 适合需要高效采样推理的场景,如自一致性验证。

我们提出一种针对多样本推理场景(如自一致性、Best-of-N采样)的新颖推测解码方法。该方法利用并行生成路径中的内在一致性,无需辅助模型或外部数据库即可合成高质量草稿令牌。通过概率聚合机制动态分析多路径间的结构模式,识别出与解码分布一致的共识令牌序列。在数学推理基准上的评估表明,相比基线方法,该方法显著提升了草稿接受率,同时降低了草稿令牌构建的延迟。本工作为高效多样本推理建立了范式转变,实现了推测解码与基于采样的推理技术的无缝融合。

原文摘要 · Abstract (English)

We propose a novel speculative decoding method tailored for multi-sample reasoning scenarios, such as self-consistency and Best-of-N sampling. Our method exploits the intrinsic consensus of parallel generation paths to synthesize high-quality draft tokens without requiring auxiliary models or external databases. By dynamically analyzing structural patterns across parallel reasoning paths through a probabilistic aggregation mechanism, it identifies consensus token sequences that align with the decoding distribution. Evaluations on mathematical reasoning benchmarks demonstrate a substantial improvement in draft acceptance rates over baselines, while reducing the latency in draft token construction. This work establishes a paradigm shift for efficient multi-sample inference, enabling seamless integration of speculative decoding with sampling-based reasoning techniques.

推理加速推测解码多样本推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。