发现大模型生成内容可能依赖记忆而非检索,提出新方法检测这一盲区。
The Attribution Blind Spot: Detecting When Language Models Rely on Memory Rather Than Retrieved Context

- 通过对比有无上下文时的内部表征差异,检测模型是否依赖记忆
- 在九种模型中验证了内部表征存在可测的轨迹特征差异
- 适合关注模型可信性与证据溯源的研究者和开发者
检索增强生成有望使大模型输出基于外部证据,但目前缺乏可靠方法验证所生成内容是否真正由检索到的上下文驱动——这对高风险应用至关重要。标准假设‘输出与上下文一致即受其控制’在检索文档与预训练数据重叠时失效:模型可能完全基于参数化记忆生成看似一致的内容,两种路径产生的输出无法区分。我们称此现象为‘归属盲区’,并提出计算现实监测(CRM)来应对。CRM借鉴认知科学中的现实监测原理,通过比较有无上下文时的内部表示,揭示仅在上下文影响下才出现的表征差异,而这类差异在输出层面被系统性忽略。CRM不判断单次生成的具体来源,而是检测预训练暴露是否在内部轨迹中留下可测量信号,构成归属推断的必要基础。在三种模型族共九个变体上,该差异集中于特定层级模式,块级噪声干预支持其一致性,且跨任务、数据集泛化良好,但在领域混淆基准上消失。归属盲区可测量且部分可解:内部表征包含输出层面不可见的诊断信号,为内知证据来源、外显可信行为的系统奠定基础。
原文摘要 · Abstract (English)
Retrieval-augmented generation promises to ground language model outputs in external evidence, yet the field has no reliable way to verify whether retrieved context actually governs generation -- a prerequisite for any high-stakes deployment. The standard assumption, that context-consistent output implies context-governed output, breaks when the retrieved document overlaps with the model's pretraining data: the model can produce faithful-looking text entirely from parametric memory, and both pathways yield indistinguishable output. We name this failure the attribution blind spot and introduce Computational Reality Monitoring (CRM) to address it. CRM operationalizes a principle adapted from cognitive science's reality monitoring framework: comparing internal representations with and without context reveals membership-conditioned representational divergence that output-level monitors systematically miss. CRM does not certify which source an individual generation used; it detects whether pretraining exposure leaves a measurable internal trajectory signature, establishing a necessary substrate for source attribution. Across nine model variants spanning three families, this divergence concentrates in architecture-specific layer patterns, receives converging support from block-level noise intervention, and generalizes across tasks and datasets while collapsing on domain-confounded benchmarks. The attribution blind spot is measurable and partially addressable: internal representations carry a diagnostic signal invisible at the output level, establishing a foundation for systems whose internal awareness of evidence provenance governs their external behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。