arXiv:2411.06528cs.CLcs.AI2024-11被引 11

发现大模型常高傲自信说错话,提出新方法让其自知不足。

Epistemic Integrity in Large Language Models

  • 构建人类标注数据集,量化模型语言自信与真实准确性的偏差。
  • 新方法使错误率比旧基准降低50%以上,验证了严重误判现象。
  • 适合关注AI可信度、安全性和可解释性的研究者与开发者。

大型语言模型日益被用作信息来源,但其以高置信度生成虚假或误导性陈述的倾向,对用户和社会构成风险。本文聚焦于认知校准失准问题——模型的语言坚定程度与其真实内部确信度不匹配。我们引入一个全新的人类标注数据集,并提出一种测量大语言模型(LLM)语言自信的新方法,相较于先前基准,错误率降低超过50%。在多个数据集上的验证表明,模型语言表达的自信与其实际准确性之间存在显著偏差。进一步的人类评估证实了这种校准失准的严重性。这一证据凸显了大模型过度自信可能带来的大规模误导风险。我们的框架为诊断此类失准提供了关键进展,也为提升跨领域AI的可信性指明了方向。

原文摘要 · Abstract (English)

Large language models are increasingly relied upon as sources of information, but their propensity for generating false or misleading statements with high confidence poses risks for users and society. In this paper, we confront the critical problem of epistemic miscalibration $\unicode{x2013}$ where a model's linguistic assertiveness fails to reflect its true internal certainty. We introduce a new human-labeled dataset and a novel method for measuring the linguistic assertiveness of Large Language Models (LLMs) which cuts error rates by over 50% relative to previous benchmarks. Validated across multiple datasets, our method reveals a stark misalignment between how confidently models linguistically present information and their actual accuracy. Further human evaluations confirm the severity of this miscalibration. This evidence underscores the urgent risk of the overstated certainty LLMs hold which may mislead users on a massive scale. Our framework provides a crucial step forward in diagnosing this miscalibration, offering a path towards correcting it and more trustworthy AI across domains.

大模型可信度校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。