arXiv:2510.13836cs.CLcs.AI2025-10EMNLP被引 1

通过生成结果的相似性,提升大模型不确定性估计的准确性。

SIMBA UQ: Similarity-Based Aggregation for Uncertainty Quantification in Large Language Models

  • 基于多个生成结果间的相似性评估置信度,无需模型内部信息。
  • 在问答、摘要等任务上,置信度校准效果优于现有方法。
  • 适合希望低成本提升大模型可信度的开发者与应用者。

大语言模型何时知道自己不知道?不确定性量化(UQ)通过估计生成结果的置信度,成为可信赖AI系统的关键组成部分。黑盒UQ方法无需访问模型内部信息,具有对系统变更鲁棒、适配不同模型、成本低等优势。本文研究以生成结果间一致性为置信度代理的UQ技术,提出一个高层非显式相似性聚合框架,涵盖多种复杂生成任务的UQ方法,并引入基于小训练集的新型置信度估计技术。在问答、摘要和文本转SQL等多样任务上的实证研究表明,所提相似性方法能获得比基线更优的校准置信度。

原文摘要 · Abstract (English)

When does a large language model (LLM) know what it does not know? Uncertainty quantification (UQ) provides measures of uncertainty, such as an estimate of the confidence in an LLM's generated output, and is therefore increasingly recognized as a crucial component of trusted AI systems. Black-box UQ methods do not require access to internal model information from the generating LLM and therefore have numerous real-world advantages, such as robustness to system changes, adaptability to choice of LLM, reduced costs, and computational tractability. In this paper, we investigate the effectiveness of UQ techniques that are primarily but not necessarily entirely black-box, where the consistency between a generated output and other sampled generations is used as a proxy for confidence in its correctness. We propose a high-level non-verbalized similarity-based aggregation framework that subsumes a broad swath of UQ approaches suitable for complex generative tasks, as well as introduce specific novel techniques from the framework that train confidence estimation models using small training sets. Through an empirical study with datasets spanning the diverse tasks of question answering, summarization, and text-to-SQL, we demonstrate that our proposed similarity-based methods can yield better calibrated confidences than baselines.

不确定性量化大模型置信度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。