不同表述方式影响大模型对用户信念的判断能力
Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It
- 通过18种表达信念的动词测试,发现模型对错误信息的信任度随动词变化
- 在'我隐约记得'场景下模型准确率比事实高50%,'我严重怀疑'则低14%
- 模型易因默认核查事实而忽略用户信念,适合需要理解用户立场的研究者
人类在日常交流中自然表达信念,如“我认为答案是3”或“我想这可能是对的”。这类信念常与事实交织,使大语言模型(LLMs)具备同时处理信念与事实的能力变得重要,尤其在面向用户的场景中。先前研究发现,即使能力强的模型也存在系统性缺陷:无法正确响应基于错误信息的用户信念。我们扩展评估至10个LLMs和18种认知表达,发现该缺陷的程度和方向取决于表达信念所用动词,对错误信息的准确率差异在‘我隐约记得’时高达+50%,而在‘我严重怀疑’时为-14%。我们进一步证明,该现象源于任务混淆:模型默认对主张进行事实核查,从而覆盖用户所表达的信念。实证显示,显式进行事实核查的思维链在错误信息上的表现反而更差;一条简单指令即可逆转不同动词族间的失败模式。机制上,模型更关注无法确认的错误信念,但解码时抑制注意力仅部分恢复准确率,且仅限于部分模型,提示需发展新的干预方法。研究澄清了既有结果,并揭示事实核查这一理想行为可能干扰信念追踪。
原文摘要 · Abstract (English)
Humans naturally form and express beliefs in daily communication, e.g., "I think the answer is 3" or "I suppose that's right." Such beliefs inevitably intertwine with fact and knowledge, making the ability to handle them in tandem desirable for large language models (LLMs), as they are increasingly deployed in user-facing settings. Prior work showed that even capable LLMs exhibit a systemic weakness in acknowledging user beliefs grounded in incorrect information. We extend this evaluation to 10 LLMs across 18 epistemic expressions and find that the size and direction of this weakness depend on the verb used to express the belief, with the accuracy gap between factual and false information ranging from +50% on "I vaguely remember" to -14% on "I seriously doubt". We further show that the phenomenon stems from what we call task confusion: models default to fact-checking the underlying claim, overriding the user's stated belief. We provide evidence where chains of thought that explicitly fact-check show lower accuracy on false information than those that do not, and a single instruction can reverse the failure across verb families. Mechanistically, models attend more to false beliefs they fail to confirm, but suppressing this attention at decoding time recovers accuracy only partially and only in some models, calling for future work on intervention methods. Our findings clarify prior results and show how fact-checking, a generally desirable behavior, can interfere with belief tracking in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。