arXiv:2601.15690cs.AIstat.AP2026-01ACL

让不确定性从诊断工具变成控制信号,提升大模型可靠性

From Passive Metric to Active Signal: The Evolving Role of Uncertainty Quantification in Large Language Models

  • 将不确定性作为主动调控信号,动态优化模型行为
  • 在推理、智能体决策和强化学习中实现自我修正与优化
  • 适合关注可信AI、模型可靠性与自主系统的研究者

尽管大语言模型展现出强大能力,其不可靠性仍是高风险领域部署的关键障碍。本文梳理了不确定性角色的演变:从被动诊断指标转向主动控制信号,以实时调控模型行为。我们展示了不确定性在三个前沿场景中的应用:在高级推理中优化计算并触发自我修正;在自主智能体中指导元认知决策,如工具使用与信息检索;在强化学习中缓解奖励劫持,通过内在奖励实现自我改进。这些进展基于贝叶斯方法和符合性预测等新兴理论框架,提供了一个统一视角。本综述系统总结了关键设计模式,强调掌握不确定性的新范式对构建可扩展、可靠且可信的下一代AI至关重要。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) show remarkable capabilities, their unreliability remains a critical barrier to deployment in high-stakes domains. This survey charts a functional evolution in addressing this challenge: the evolution of uncertainty from a passive diagnostic metric to an active control signal guiding real-time model behavior. We demonstrate how uncertainty is leveraged as an active control signal across three frontiers: in \textbf{advanced reasoning} to optimize computation and trigger self-correction; in \textbf{autonomous agents} to govern metacognitive decisions about tool use and information seeking; and in \textbf{reinforcement learning} to mitigate reward hacking and enable self-improvement via intrinsic rewards. By grounding these advancements in emerging theoretical frameworks like Bayesian methods and Conformal Prediction, we provide a unified perspective on this transformative trend. This survey provides a comprehensive overview, critical analysis, and practical design patterns, arguing that mastering the new trend of uncertainty is essential for building the next generation of scalable, reliable, and trustworthy AI.

大模型不确定性可信AI智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。