arXiv:2506.07448cs.LGcs.AI2025-06被引 1

用贝叶斯实验建模提升大模型不确定性管理能力

Extending Epistemic Uncertainty Beyond Parameters Would Assist in Designing Reliable LLMs

  • 引入贝叶斯实验建模框架,系统区分不同来源的不确定性
  • 支持主动纠错而非简单拒绝高不确定输出
  • 适合高风险场景下需可靠推理的LLM应用

尽管大语言模型(LLMs)具有高度交互性和可扩展性,但当前确保部署可靠性的方式仍主要局限于对高不确定性输出进行拒绝,以避免错误信息。这种保守策略反映出缺乏系统工具来区分和应对不同类型的不确定性。本文倡导采用贝叶斯实验建模(Bayesian Modeling of Experiments)框架,该框架为推理不确定性提供了连贯基础,并明确区分了可减少与不可减少的不确定性。该框架使大模型及其用户能够根据上下文采取适当行动,如请求澄清、检索外部信息或优化输入。通过支持主动解决而非被动回避,为更可靠、透明且广泛适用的大模型系统开辟了道路,尤其适用于高风险的真实世界场景。

原文摘要 · Abstract (English)

Although large language models (LLMs) are highly interactive and extendable, current approaches to ensure reliability in deployments remain mostly limited to rejecting outputs with high uncertainty in order to avoid misinformation. This conservative strategy reflects the current lack of tools to systematically distinguish and respond to different sources of uncertainty. In this paper, we advocate for the adoption of Bayesian Modeling of Experiments -- a framework that provides a coherent foundation to reason about uncertainty and clarify the reducibility of uncertainty -- for managing and proactively addressing uncertainty that arises in LLM deployments. This framework enables LLMs and their users to take contextually appropriate steps, such as requesting clarification, retrieving external information, or refining inputs. By supporting active resolution rather than passive avoidance, it opens the door to more reliable, transparent, and broadly applicable LLM systems, particularly in high-stakes, real-world settings.

大模型可靠性不确定性建模贝叶斯方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。