arXiv:2506.21849cs.CLcs.AI2025-06中稿 · The Conference on …被引 7

用生成一致性估算大模型置信度,实证其有效性并提出新方法。

The Consistency Hypothesis in Uncertainty Quantification for Large Language Models

  • 以多次生成结果的相似性作为置信度代理,验证其合理性。
  • 在8个数据集3类任务上验证假设普遍成立,最有效的是'任意相似性'假设。
  • 提出无需数据的黑箱方法,通过生成间相似度聚合提升置信度估计性能。

估算大语言模型输出的置信度对需要高用户信任的真实应用至关重要。依赖模型API访问的黑箱不确定性量化(UQ)方法因其实用性而受到欢迎。本文研究了若干UQ方法背后的隐含假设——以生成一致性作为置信度的代理,我们将其形式化为一致性假设。提出了三个数学表述及相应的统计检验,用于捕捉该假设的不同变体,并设计指标评估模型在不同任务中的输出一致性。我们的实证研究覆盖8个基准数据集和3项任务(问答、文本摘要、文本转SQL),揭示了该假设在不同设置下的普遍性。其中,'Sim-Any'假设最具可操作性,我们据此提出无需数据的黑箱UQ方法,通过聚合多次生成间的相似性来估计置信度。该方法优于现有基线,展示了实证发现的一致性假设的实际价值。

原文摘要 · Abstract (English)

Estimating the confidence of large language model (LLM) outputs is essential for real-world applications requiring high user trust. Black-box uncertainty quantification (UQ) methods, relying solely on model API access, have gained popularity due to their practical benefits. In this paper, we examine the implicit assumption behind several UQ methods, which use generation consistency as a proxy for confidence, an idea we formalize as the consistency hypothesis. We introduce three mathematical statements with corresponding statistical tests to capture variations of this hypothesis and metrics to evaluate LLM output conformity across tasks. Our empirical investigation, spanning 8 benchmark datasets and 3 tasks (question answering, text summarization, and text-to-SQL), highlights the prevalence of the hypothesis under different settings. Among the statements, we highlight the `Sim-Any' hypothesis as the most actionable, and demonstrate how it can be leveraged by proposing data-free black-box UQ methods that aggregate similarities between generations for confidence estimation. These approaches can outperform the closest baselines, showcasing the practical value of the empirically observed consistency hypothesis.

大模型置信度黑箱一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。