arXiv:2410.14582cs.AIcs.CL2024-10ICLR被引 29

首次系统评估大模型指令遵循中的不确定性估计能力

Do LLMs estimate uncertainty well in instruction-following?

  • 设计受控实验环境,分离指令遵循误差与不确定性来源
  • 现有方法在细微错误场景下表现差,内部状态改进有限
  • 为高风险应用中可信AI代理提供关键参考

大型语言模型(LLMs)若能精准遵循用户指令,可成为各领域有价值的个人智能代理。然而,近期研究揭示了其在指令遵循方面存在显著局限,引发对高风险应用场景可靠性的担忧。准确估计模型在遵循指令时的不确定性,对降低部署风险至关重要。本文首次系统评估了LLMs在指令遵循任务中的不确定性估计能力。研究发现,现有指令遵循基准存在多重因素混杂问题,难以分离不确定性来源,阻碍方法与模型间的比较。为此,我们提出两种受控数据集版本,支持在不同条件下全面对比不确定性估计方法。结果表明,现有方法在模型出现细微指令遵循错误时表现不佳;尽管利用内部模型状态有一定改善,但在复杂场景下仍不充分。本研究的受控评估框架揭示了大模型在不确定性估计方面的局限性与潜力,为构建更可信的AI代理提供了重要基础。

原文摘要 · Abstract (English)

Large language models (LLMs) could be valuable personal AI agents across various domains, provided they can precisely follow user instructions. However, recent studies have shown significant limitations in LLMs' instruction-following capabilities, raising concerns about their reliability in high-stakes applications. Accurately estimating LLMs' uncertainty in adhering to instructions is critical to mitigating deployment risks. We present, to our knowledge, the first systematic evaluation of the uncertainty estimation abilities of LLMs in the context of instruction-following. Our study identifies key challenges with existing instruction-following benchmarks, where multiple factors are entangled with uncertainty stems from instruction-following, complicating the isolation and comparison across methods and models. To address these issues, we introduce a controlled evaluation setup with two benchmark versions of data, enabling a comprehensive comparison of uncertainty estimation methods under various conditions. Our findings show that existing uncertainty methods struggle, particularly when models make subtle errors in instruction following. While internal model states provide some improvement, they remain inadequate in more complex scenarios. The insights from our controlled evaluation setups provide a crucial understanding of LLMs' limitations and potential for uncertainty estimation in instruction-following tasks, paving the way for more trustworthy AI agents.

大模型不确定性指令遵循

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。