arXiv:2606.28327cs.IRcs.AI2026-06中稿 · CogSci 2026

对比人脑与AI系统在语义干扰下的检索能力,发现人脑更抗干扰。

The Interference Gap: Comparing Retrieval Bounds in Human Memory and RAG Systems

  • 用信号检测理论统一建模人脑与RAG系统的检索过程
  • 人脑干扰敏感度(α/σ=0.41)低于密集检索的RAG系统(0.67)
  • 提出可验证的假设,连接认知科学与AI评估

在语义干扰条件下,人类情景记忆与检索增强生成(RAG)系统的检索边界有何差异?我们提出一个统一的信号检测理论(SDT)框架,适用于两者,并在匹配实验范式中拟合行为与计算数据。两种系统均表现出准确率随关联数量(扇出)对数下降,但人类干扰敏感度更低(α/σ=0.41),低于密集段落检索的RAG系统(α/σ=0.67),而受认知启发的HippoRAG介于二者之间(α/σ=0.44)。行为实验(N=112)与模拟验证了该框架;参数恢复显示可识别性高(r ≥ .93),模型比较支持对数形式优于幂律替代模型(ΔBIC > 15)。我们讨论编码特异性、时间上下文绑定和检索门控作为候选机制,其因果作用有待验证。六个可证伪预测将认知记忆研究与AI检索评估联系起来。

原文摘要 · Abstract (English)

How do retrieval bounds compare between human episodic memory and Retrieval-Augmented Generation (RAG) systems under semantic interference? We present a unified signal detection theory (SDT) framework that applies to both, and use it to fit behavioral and computational data in matched paradigms. Both systems show logarithmic accuracy decline with association count (fan), but humans exhibit lower interference sensitivity ($α/σ= 0.41$) than dense passage retrieval ($α/σ= 0.67$), with cognitively-inspired HippoRAG falling between the two ($α/σ= 0.44$). Behavioral experiments ($N = 112$) and simulations validate the framework; parameter recovery confirms identifiability ($r \geq .93$) and model comparison favors the logarithmic specification over a power-law alternative ($Δ$BIC $> 15$). We discuss encoding specificity, temporal context binding, and retrieval gating as candidate mechanisms whose causal role remains to be established. Six falsifiable predictions connect cognitive memory research with AI retrieval evaluation.

记忆模型RAG系统认知科学信号检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。