arXiv:2604.10981cs.AIcs.IR2026-04被引 1

厘清大模型记忆连续性评估的盲区,指出现有基准均无法完整衡量连续性。

ATANT v1.1: Positioning Continuity Evaluation Against Memory, Long-Context, and Agentic-Memory Benchmarks

  • 通过结构分析对比7项关键属性,发现现有评测平均仅覆盖0.43项
  • 实测显示其LOCOMO评分8.8%与自身96%连续性得分差异巨大
  • 提醒研究者避免混淆不同评估体系,专注真正定义连续性的能力

ATANT v1.0(arXiv:2604.06710)将连续性定义为系统属性,并提出包含7个必要特性的评估框架,采用无需LLM的10检查点方法,在250篇故事语料上验证。本补充论文v1.1不修改原标准,而是填补原版因页数限制未充分讨论的相关工作空白。通过结构分析表明,现有基准如LOCOMO、LongMemEval、BEAM、MemoryBench、Zep评估套件、Letta/MemGPT评测及RULER均未完整覆盖该连续性定义的7项属性:中位数覆盖1项,均值仅0.43项(部分计分按0.5折算),无一超过2项。本文提供逐项属性覆盖矩阵,揭示各基准方法学缺陷(包括LOCOMO参考实现中的空真值打分漏洞,导致其23%语料本不可评分),并公开自身在LOCOMO上的得分为8.8%,说明该结果对连续性无意义。将此与自身的96% ATANT累积尺度得分对照,87点差距证明两者测量不同特性,非性能优劣之别。立场非对抗性:各基准衡量真实能力,但均无法裁定连续性,领域因此长期忽视v1.0所列属性。

原文摘要 · Abstract (English)

ATANT v1.0 (arXiv:2604.06710) defined continuity as a system property with 7 required properties and introduced a 10-checkpoint, LLM-free evaluation methodology validated on a 250-story corpus. Since publication, a recurring reviewer and practitioner question has concerned not the framework itself but its relationship to a wider set of memory evaluations: LOCOMO, LongMemEval, BEAM, MemoryBench, Zep's evaluation suite, Letta/MemGPT's evaluations, and RULER. This companion paper, v1.1, does not modify the v1.0 standard. It closes a related-work gap that v1.0 left brief under page limits. We show by structural analysis that none of these benchmarks measures continuity as defined in v1.0: of the 7 required properties, the median existing eval covers 1 property, the mean covers 0.43 when partial credit is scored at 0.5, and no eval covers more than 2. We provide a cell-by-cell property-coverage matrix, identify methodological defects specific to each benchmark (including an empty-gold scoring bug in the LOCOMO reference implementation that renders 23% of its corpus unscorable by construction), and publish our reference implementation's LOCOMO score (8.8%) alongside the structural reason that number is uninformative about continuity. We publish our 8.8% LOCOMO score alongside our 96% ATANT cumulative-scale score as a calibration pair: the 87-point divergence is evidence that the two benchmarks measure different properties, not that one system is an order of magnitude better than another. The position v1.1 takes is not adversarial: each benchmark measures a real capability. The claim is that none of them can adjudicate continuity, and conflating them with continuity evaluation has led the field to under-invest in the properties v1.0 names.

大模型评估记忆连续性基准测试框架对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。