arXiv:2410.20199cs.AI2024-10综述被引 6

厘清大模型预测不确定性来源,提升可信度

Rethinking the Uncertainty: A Critical Review and Analysis in the Era of Large Language Models

  • 构建大模型不确定性分类框架,系统识别来源
  • 指出现有方法仅关注置信度,忽略真实不确定性
  • 适合安全关键场景研究者与模型可靠性开发者

近年来,大语言模型(LLMs)已成为人工智能应用的核心。随着其应用范围扩大,准确估计预测不确定性变得至关重要。当前方法往往难以精确识别、衡量和应对真实不确定性,多数仅聚焦于模型置信度的估算。这一差距主要源于对不确定性注入位置、时机和方式的理解不完整。本文提出一个专门针对大模型特性的综合框架,旨在识别并理解不确定性类型与来源。该框架通过系统分类与定义各类不确定性,深化了对复杂不确定性格局的认识,为开发精准量化方法奠定基础。同时,文章详细阐述关键概念,分析现有方法在任务关键与安全敏感场景中的局限性,并展望未来方向,以提升这些方法在真实世界中的可靠性和可采纳性。

原文摘要 · Abstract (English)

In recent years, Large Language Models (LLMs) have become fundamental to a broad spectrum of artificial intelligence applications. As the use of LLMs expands, precisely estimating the uncertainty in their predictions has become crucial. Current methods often struggle to accurately identify, measure, and address the true uncertainty, with many focusing primarily on estimating model confidence. This discrepancy is largely due to an incomplete understanding of where, when, and how uncertainties are injected into models. This paper introduces a comprehensive framework specifically designed to identify and understand the types and sources of uncertainty, aligned with the unique characteristics of LLMs. Our framework enhances the understanding of the diverse landscape of uncertainties by systematically categorizing and defining each type, establishing a solid foundation for developing targeted methods that can precisely quantify these uncertainties. We also provide a detailed introduction to key related concepts and examine the limitations of current methods in mission-critical and safety-sensitive applications. The paper concludes with a perspective on future directions aimed at enhancing the reliability and practical adoption of these methods in real-world scenarios.

大模型不确定性可信AI可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。