arXiv:2512.09538stat.MLcs.CL2025-12被引 2

用束搜索提升大模型不确定性估计的稳定性与准确性

Don't Throw Away Your Beams: Improving Consistency-based Uncertainties in LLMs via Beam Search

  • 改用束搜索生成候选答案,替代随机采样
  • 在6个问答数据集上显著降低估计方差并提升性能
  • 理论证明束搜索在特定条件下优于传统采样方法

一致性方法已成为大语言模型不确定性量化(UQ)的有效手段。这类方法通常依赖多个通过多项式采样生成的输出,通过测量其一致程度来评估不确定性。然而,在短文本问答任务中,多项式采样易因分布尖锐产生重复结果,且其随机性导致不同运行间的不确定性估计波动较大。本文提出一种新方法,采用束搜索生成一致性检验的候选输出,相比多项式采样显著提升了性能并降低了估计方差。我们还推导了束搜索集合概率质量的理论下界,证明其在该条件下误差小于多项式采样。在六个问答数据集上的实验表明,该方法持续优于现有方法,达到当前最优的不确定性量化表现。

原文摘要 · Abstract (English)

Consistency-based methods have emerged as an effective approach to uncertainty quantification (UQ) in large language models. These methods typically rely on several generations obtained via multinomial sampling, measuring their agreement level. However, in short-form QA, multinomial sampling is prone to producing duplicates due to peaked distributions, and its stochasticity introduces considerable variance in uncertainty estimates across runs. We introduce a new family of methods that employ beam search to generate candidates for consistency-based UQ, yielding improved performance and reduced variance compared to multinomial sampling. We also provide a theoretical lower bound on the beam set probability mass under which beam search achieves a smaller error than multinomial sampling. We empirically evaluate our approach on six QA datasets and find that its consistent improvements over multinomial sampling lead to state-of-the-art UQ performance.

大模型不确定性束搜索问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。