arXiv:2505.22655cs.LGcs.AI2025-05ICML被引 32

LLM代理交互中需重新思考不确定性量化方式

Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents

  • 提出三类新不确定性:任务不明确、交互式学习、输出表达
  • 传统不确定性分类在交互场景中失效且自相矛盾
  • 适合关注AI可信性与人机交互的 researchers

大型语言模型(LLMs)和聊天机器人代理时常产生错误输出,且此类错误无法完全避免。因此,不确定性量化至关重要,旨在以单一数值或分别量化认知型与随机型不确定性。本文指出,这种传统二分法在开放、互动的LLM代理使用场景中过于局限,亟需拓展。我们回顾文献发现,现有对认知型与随机型不确定性的定义相互矛盾,在交互式设置中失去意义。为此,我们提出三个新研究方向:1)未明确定义的不确定性(当用户未完整提供信息或任务边界不清时);2)交互式学习(通过追问减少上下文不确定性);3)输出不确定性(利用丰富语言与语音空间表达不确定性,而非仅依赖数值)。这些新方法有望使LLM代理交互更透明、可信赖、直观。

原文摘要 · Abstract (English)

Large-language models (LLMs) and chatbot agents are known to provide wrong outputs at times, and it was recently found that this can never be fully prevented. Hence, uncertainty quantification plays a crucial role, aiming to quantify the level of ambiguity in either one overall number or two numbers for aleatoric and epistemic uncertainty. This position paper argues that this traditional dichotomy of uncertainties is too limited for the open and interactive setup that LLM agents operate in when communicating with a user, and that we need to research avenues that enrich uncertainties in this novel scenario. We review the literature and find that popular definitions of aleatoric and epistemic uncertainties directly contradict each other and lose their meaning in interactive LLM agent settings. Hence, we propose three novel research directions that focus on uncertainties in such human-computer interactions: Underspecification uncertainties, for when users do not provide all information or define the exact task at the first go, interactive learning, to ask follow-up questions and reduce the uncertainty about the current context, and output uncertainties, to utilize the rich language and speech space to express uncertainties as more than mere numbers. We expect that these new ways of dealing with and communicating uncertainties will lead to LLM agent interactions that are more transparent, trustworthy, and intuitive.

不确定性量化人机交互LLM代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。