arXiv:2505.07309cs.LG2025-05被引 3

分解大模型不确定性来源,实现精准选型与评估。

Uncertainty Profiles for LLMs: Uncertainty Source Decomposition and Adaptive Model-Metric Selection

  • 将不确定性拆分为四类,构建针对性估计流程。
  • 不同任务和模型的不确定性特征差异显著。
  • 根据任务特性动态选择模型或度量指标。

大型语言模型常生成流畅但事实错误的输出,即幻觉,影响其在实际应用中的可靠性。尽管不确定性估计被视为检测此类错误的有前景策略,但现有度量方法可解释性差,且难以明确其捕捉的不确定性类型。本文提出一种系统性框架,将大模型不确定性分解为四种独立来源,借鉴前期研究设计源特定估计流程,量化各类不确定性,并评估现有度量在不同任务与模型中对各来源的响应。结果表明,度量、任务与模型在不确定性特征上存在系统性差异。基于此,我们提出一种由任务不确定性特征引导的模型/度量自适应选择方法。跨数据集与模型的实验表明,该不确定性感知选择策略持续优于基线,有助于精准选择合适模型或度量,提升不确定性估计的可靠性与部署效率。

原文摘要 · Abstract (English)

Large language models (LLMs) often generate fluent but factually incorrect outputs, known as hallucinations, which undermine their reliability in real-world applications. While uncertainty estimation has emerged as a promising strategy for detecting such errors, current metrics offer limited interpretability and lack clarity about the types of uncertainty they capture. In this paper, we present a systematic framework for decomposing LLM uncertainty into four distinct sources, inspired by previous research. We develop a source-specific estimation pipeline to quantify these uncertainty types and evaluate how existing metrics relate to each source across tasks and models. Our results show that metrics, task, and model exhibit systematic variation in uncertainty characteristic. Building on this, we propose a method for task specific metric/model selection guided by the alignment or divergence between their uncertainty characteristics and that of a given task. Our experiments across datasets and models demonstrate that our uncertainty-aware selection strategy consistently outperforms baseline strategies, helping us select appropriate models or uncertainty metrics, and contributing to more reliable and efficient deployment in uncertainty estimation.

大模型不确定性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。