arXiv:2603.07887cs.LGcs.AI2026-03被引 1

用粒子滤波理论解析大模型推理中多样本采样的准确率与成本关系

Reject, Resample, Repeat: Understanding Parallel Reasoning in Language Model Inference

  • 将语言模型推理视为粒子滤波过程,构建可证明的采样框架
  • 提出非渐近保证条件,验证采样误差与评估次数的定量关系
  • 揭示所有此类方法的根本限制,为优化提供理论依据

推理阶段通过聚合和修剪多个样本已成为引导大型语言模型的重要范式,但对其准确率与计算成本之间权衡关系缺乏系统理解。本文引入粒子滤波算法(如序列蒙特卡洛,SMC)作为分析工具,基于基础语言模型和一个过程奖励模型(估计最终奖励),研究在给定过程奖励评估次数的前提下,能以多高精度从目标分布中采样。理论上,我们识别出(1)实现非渐近保证的简单条件;(2)对SMC的算法改进;(3)所有粒子滤波方法面临的根本极限。实验上,我们的理论条件有效控制了SMC的采样误差,但未必决定最终准确性,暗示需发展超越采样的理论视角。

原文摘要 · Abstract (English)

Inference-time methods that aggregate and prune multiple samples have emerged as a powerful paradigm for steering large language models, yet we lack any principled understanding of their accuracy-cost tradeoffs. In this paper, we introduce a route to rigorously study such approaches using the lens of *particle filtering* algorithms such as Sequential Monte Carlo (SMC). Given a base language model and a *process reward model* estimating expected terminal rewards, we ask: *how accurately can we sample from a target distribution given some number of process reward evaluations?* Theoretically, we identify (1) simple criteria enabling non-asymptotic guarantees for SMC; (2) algorithmic improvements to SMC; and (3) a fundamental limit faced by all particle filtering methods. Empirically, we demonstrate that our theoretical criteria effectively govern the *sampling error* of SMC, though not necessarily its final *accuracy*, suggesting that theoretical perspectives beyond sampling may be necessary.

语言模型推理优化粒子滤波采样分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。