arXiv:2504.19622cs.AI2025-04NAACL被引 3

用贝叶斯认识论分析语言模型如何根据证据调整信心与回答。

From Evidence to Belief: A Bayesian Epistemology Approach to Language Models

  • 从贝叶斯认识论出发,考察模型对不同可靠性和信息量证据的响应机制。
  • 模型在面对真实证据时符合贝叶斯确认假设,但整体不遵循所有贝叶斯原则。
  • 模型易受黄金证据偏见影响,高信心不等于高准确率,适合研究可信度评估者阅读。

本文从贝叶斯认识论视角探究语言模型的知识表征。通过构建包含多种类型证据的数据集,分析模型在面对不同信息量与可靠性证据时的响应与置信度变化,采用口头置信度、词元概率和采样方法进行评估。结果发现,语言模型在真实证据下较好遵循贝叶斯确认假设,但在其他情况下违背贝叶斯原则。此外,模型在强证据下可能表现出高置信度,但这并不总对应高准确性。分析还揭示模型对黄金证据存在偏倚,且其表现随证据无关程度而异,解释了为何模型偏离贝叶斯假设。

原文摘要 · Abstract (English)

This paper investigates the knowledge of language models from the perspective of Bayesian epistemology. We explore how language models adjust their confidence and responses when presented with evidence with varying levels of informativeness and reliability. To study these properties, we create a dataset with various types of evidence and analyze language models' responses and confidence using verbalized confidence, token probability, and sampling. We observed that language models do not consistently follow Bayesian epistemology: language models follow the Bayesian confirmation assumption well with true evidence but fail to adhere to other Bayesian assumptions when encountering different evidence types. Also, we demonstrated that language models can exhibit high confidence when given strong evidence, but this does not always guarantee high accuracy. Our analysis also reveals that language models are biased toward golden evidence and show varying performance depending on the degree of irrelevance, helping explain why they deviate from Bayesian assumptions.

贝叶斯认知语言模型置信度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。