arXiv:2609.01852cs.AIcs.CL2026-09

大模型越强,越容易被过时记忆误导,反而不如小模型可靠。

The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents

  • 用不同规模模型测试记忆依赖性,发现大模型更易盲目信任过时信息
  • 在关键任务中,80亿参数模型在错误记忆下准确率暴跌至0.92以下
  • 适合研究大模型可靠性、长期记忆机制的开发者和研究人员

持续记忆虽支持个性化代理,但过时信息可能无预警覆盖权威证据。我们研究模型能力变化时该危害何时出现。在相同家族模型系列(Qwen3 0.6/1.7/4/8B)上,评估一个冻结的封闭集动作评分基准,包含两个套件:受益套件(无记忆无法解决)、安全套件(权威工具始终正确)。记忆信任鸿沟源于过度信任而非混淆。在受益套件中,各规模模型使用过时答案占比达0.92–1.00。在安全套件中,危害低于无记忆基线($ riangle_{ ext{mem}}$)的现象具有能力门槛,大模型一旦过时信息看似更新即迅速崩溃。四因素因子实验显示,触发过度信任的因素取决于特征与模型规模。移除标签会放大所有规模的过度信任;近期性特征(过时日期却显示较新)更易欺骗大模型。来源权威性弱且规模不变,位置从正转负贯穿Qwen3系列。通过跨规模直接对比验证这些交互作用,而非重叠区间。缓解策略也具能力依赖:暴露元数据仅提升有能力模型的准确率,而仅提前解决冲突才能恢复小型检查点准确率。该模式在独立的Llama-Instruct系列及两个外部数据集(RGB, MisBench)中重现。控制实验表明记忆标签无稳定优势:在0.6/1.7/4B,模型更信任过时文档而非过时记忆;8B时差异不显著。

原文摘要 · Abstract (English)

Persistent memory supports personalized agents, but a stale stored fact can override current authoritative evidence without warning. We study when this harm begins as model capability changes. We evaluate a frozen, closed-set, action-scored benchmark with 2 suites that represent 2 different meanings of "no memory" (a Benefit suite, unsolvable without the stored fact, and a Safety suite, in which an authoritative tool always holds the correct value), on a same-family model-size series (Qwen3 0.6/1.7/4/8B). The Memory Trust Gap reflects over-trust rather than confusion. In the Benefit suite, models answer with the stale value 0.92-1.00 of the time at every scale. In the Safety suite, harm below the no-memory baseline under the trap conditions ($\Delta_{\mathrm{mem}}$) is capability-gated, with the larger models collapsing most once a stale note is made to look current. In a $2\times2\times2\times2$ factorial, which feature triggers over-trust depends on both the feature and model scale. Removing a label amplifies over-trust at every size, and a recency feature (stale dated newer) fools the larger models harder. Source authority is weak and scale-flat, and position changes from positive to negative across the Qwen3 model-size series. We confirm these scale interactions with direct cross-size contrast tests rather than overlapping per-model intervals. Mitigation is likewise capability-dependent: exposing metadata improves accuracy for the capable models, but only pre-resolving the conflict restores accuracy for the 2 smaller checkpoints. The same pattern appears on the capable models in an independent Llama-Instruct model-size series and on 2 external datasets (RGB, MisBench). A framing control finds no consistent advantage for the memory label: at the 3 smaller scales, models trust a stale document more than a stale memory; at 8B, the difference is not significant.

大模型记忆偏差可靠性模型规模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。