同一记忆测试中,评分目标不同会导致结论翻转。
Same Ranking, Different Winner: How Scoring Targets Shape LLM Memory Benchmarks
- 提出TIAP审计方法,无需重跑检索即可重评三种评分目标
- 83.4%~94.0%查询结果因评分目标变化而改变nDCG排名
- 多数情况下宽松的源关联评分不合理,但论文常默认使用
对话记忆系统将对话历史转化为事实、摘要、时间线等与源关联的衍生记忆,导致单个源轮次可同时存在多个存储形式。这引发一个未明确定义的评估问题:应将检索得分归于哪种存储形式?我们发现,评分目标的选择常被隐含,默认设置会显著影响基准结论。本文提出TIAP——一种固定输出审计方法,可在不重新运行检索的情况下,对保存的排序结果按原始(Raw)、源(Source)和规范(Canonical)三种目标重评分。在LoCoMo和LongMemEval-S数据集上,仅改变评分目标就使83.4%至94.0%的共享查询的nDCG发生变化,导致Mem0和MemoryOS迁移实验中的目标排序反转,并改变解析器密度推荐。1,902例语义审计显示,仅29.2%情况下放宽源关联评分是合理的,尽管验证子集的评分一致性较高。结果表明,评分目标具有非不变性:记忆架构的结论可能仅因基准设计的一个选择而悄然翻转。因此,对话记忆相关研究应明确并报告其评分目标。
原文摘要 · Abstract (English)
Conversational-memory systems increasingly transform dialogue history into facts, summaries, timelines, and other source-linked descendants, so a single source turn can coexist with several derived memories in the same retrieval index. This raises an underspecified evaluation question: which stored form should receive retrieval credit? We show that this scoring-target choice is often left implicit and can materially change benchmark conclusions. We present TIAP, a fixed-output audit that rescores saved ranked outputs under three targets -- Raw, Source, and Canonical -- without rerunning retrieval. On LoCoMo and LongMemEval-S, switching only the credited target changes nDCG on 83.4--94.0 percent of shared queries, flips target orderings on Mem0 and MemoryOS transfer runs, and reverses parser-density recommendations. A 1,902-case semantic audit further shows that relaxed source-linked credit is fully justified only 29.2 percent of the time, despite high rubric reliability in a validation subset. These results reveal target noninvariance: conclusions about memory architectures can silently flip with a single benchmark-design choice. Conversational-memory papers should therefore define and report the scoring target explicitly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。