测试大模型对事实、信念和知识的区分能力,发现其在判断个人信念时严重失误。
Belief in the Machine: Investigating Epistemological Blind Spots of Language Models
- 构建新数据集KaBLE,评估大模型在13类认知任务中的表现
- 模型对第一人称信念判断准确率仅54.4%,远低于第三人称的80.7%
- 缺乏对知识必须为真的本质理解,易受语言表面线索误导
随着语言模型在医疗、法律和新闻等领域的广泛应用,其区分事实、信念与知识的能力至关重要。本文系统评估了GPT-4、Claude-3和Llama-3等现代语言模型的元认知推理能力,采用包含13,000个问题的KaBLE数据集,覆盖13项任务。结果显示:模型在事实场景中准确率达86%,但在虚假场景下表现显著下降;尤其在个人信念判断上存在明显缺陷,当信念与事实冲突时准确率大幅降低;第一人称信念任务准确率为54.4%,远低于第三人称的80.7%;模型未掌握知识的真值性(factive nature),常依赖语言线索而非深层推理进行事实判断。这些发现揭示当前模型在关键认知维度上的重大局限,警示其在医疗咨询、心理支持等高风险场景中的部署风险。
原文摘要 · Abstract (English)
As language models (LMs) become integral to fields like healthcare, law, and journalism, their ability to differentiate between fact, belief, and knowledge is essential for reliable decision-making. Failure to grasp these distinctions can lead to significant consequences in areas such as medical diagnosis, legal judgments, and dissemination of fake news. Despite this, current literature has largely focused on more complex issues such as theory of mind, overlooking more fundamental epistemic challenges. This study systematically evaluates the epistemic reasoning capabilities of modern LMs, including GPT-4, Claude-3, and Llama-3, using a new dataset, KaBLE, consisting of 13,000 questions across 13 tasks. Our results reveal key limitations. First, while LMs achieve 86% accuracy on factual scenarios, their performance drops significantly with false scenarios, particularly in belief-related tasks. Second, LMs struggle with recognizing and affirming personal beliefs, especially when those beliefs contradict factual data, which raises concerns for applications in healthcare and counseling, where engaging with a person's beliefs is critical. Third, we identify a salient bias in how LMs process first-person versus third-person beliefs, performing better on third-person tasks (80.7%) compared to first-person tasks (54.4%). Fourth, LMs lack a robust understanding of the factive nature of knowledge, namely, that knowledge inherently requires truth. Fifth, LMs rely on linguistic cues for fact-checking and sometimes bypass the deeper reasoning. These findings highlight significant concerns about current LMs' ability to reason about truth, belief, and knowledge while emphasizing the need for advancements in these areas before broad deployment in critical sectors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。