让大模型说真话:用自然语言准确表达不确定性的方法
MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
- 基于人类元认知设计新提示策略,让模型更真实地表达不确定性
- 在多项任务中提升61%的表达真实性,83%的人类判断胜过原始输出
- 发现现有校准方法反而损害真实性,强调语言表达需与内在信心一致
大模型可信度的关键在于可靠地传达不确定性,但当前模型常以肯定语气陈述错误信息,导致用户过度依赖并降低信任。本文首次系统研究大模型的'忠实置信度校准'能力,评估不同模型、数据集和提示策略下,模型使用语言表达不确定性是否真实反映其内在置信度。结果表明,大模型普遍表现不佳,现有干预手段效果有限:标准提示仅带来微弱提升,基于事实性的校准技术甚至可能损害忠实性。为此,我们提出MetaFaith——一种受人类元认知启发的新型提示校准方法。实验显示,MetaFaith在多种模型和任务中均显著提升忠实性,使表达真实性最高提升61%,且在人类评估中83%优于原始生成结果。
原文摘要 · Abstract (English)
A critical component in the trustworthiness of LLMs is reliable uncertainty communication, yet LLMs often use assertive language when conveying false claims, leading to over-reliance and eroded trust. We present the first systematic study of $\textit{faithful confidence calibration}$ of LLMs, benchmarking models' ability to use linguistic expressions of uncertainty that $\textit{faithfully reflect}$ their intrinsic uncertainty, across a comprehensive array of models, datasets, and prompting strategies. Our results demonstrate that LLMs largely fail at this task, and that existing interventions are insufficient: standard prompt approaches provide only marginal gains, and existing, factuality-based calibration techniques can even harm faithful calibration. To address this critical gap, we introduce MetaFaith, a novel prompt-based calibration approach inspired by human metacognition. We show that MetaFaith robustly improves faithful calibration across diverse models and task domains, enabling up to 61% improvement in faithfulness and achieving an 83% win rate over original generations as judged by humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。