arXiv:2607.23927cs.AIcs.CL2026-07

测试大模型对话中区分自我生成与外部信息的能力,发现记忆结构影响判断准确率。

Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory

论文配图:Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory
图 1 · 摘自论文原文
  • 通过控制对话记忆结构,测试模型对内容来源的判断机制。
  • 记忆延迟下,模型自产内容识别准确率从接近满分降至劣势。
  • 部分模型出现判断错位或自信与正确性脱钩,现有评估无法捕捉此问题。

对话式AI若无法区分自身输出与用户输入,会将自身错误当作用户提供的事实。在人类中,这种能力称为现实监控,其失效与幻觉、妄想和编造有关,但大语言模型是否具备该能力尚无验证。本文在两项实验及六种大模型上表明,来源归因依赖于对话记忆的构建方式:在记忆需求较小时,模型对自产内容的识别准确率接近满分;但当引入情景延迟后,该优势逆转为对外部信息的脆弱偏好。反馈显示两种缺陷:部分模型内部与外部判断发生互换;另一些模型虽准确率提升,但自信程度与正确性脱钩,此类现象未被现有基准检测到。跨模型分析表明,活跃参数量而非总参数量才是关键因素。这提示:随着AI系统承担自主多轮任务,评估其‘知道什么’已不够,还需追踪‘这些知识来自何处’。

原文摘要 · Abstract (English)

A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts. In humans, this capacity is called reality monitoring, and its failures are linked to hallucinations, delusions, and confabulation, yet whether LLMs possess it remains untested. Here we show, across two experiments and six LLMs, that source attribution depends on how conversational memory is structured: ceiling accuracy for self-generated content under minimal memory demands reverses to a fragile external-item advantage once episodic delay removes that shortcut. Feedback exposes two failures: in some models, internal and external judgments swap; in others, accuracy improves while confidence decouples from correctness, dissociations invisible to existing benchmarks. Across models, this pattern implicates active, not aggregate, parameter count. This suggests that as AI systems take on autonomous, multi-turn roles, evaluating what they know is not enough: tracking where that knowledge came from may matter equally.

大模型认知记忆机制源归属

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。