测试长时智能体对复杂记忆关系的细微区分能力
SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents

- 构建受控关系的语义记忆片段,模拟真实交互中的复杂记忆
- 1522个评测实例显示现有系统在关系辨别上仍表现薄弱
- 适合评估长期记忆系统的推理与检索能力,尤其关注关系保持
持久性AI助手(如OpenClaw)在长期交互中积累大量相关记忆。随着记忆增长,它们可能相互强化、跨场景分化或直接冲突,正确服务依赖于记忆间的关系而非孤立回忆。现有长期记忆基准很少检验智能体在下游任务中如何保留和利用这些关系。为此,我们提出SubtleMemory,一个针对长周期智能体细粒度关系记忆区分的基准。该基准构造受关系控制的潜在语义特征,其变体体现互补、微妙或矛盾关系,并嵌入真实的用户-代理历史中,要求智能体在后续查询和指令中恢复分布式的关联结构。基准包含1,522个评估实例,覆盖10条长历史,基于1,090组受控关系的记忆变体,涵盖用户相关与非用户相关查询。评估六种独立记忆系统、两种带原生记忆模块的Claw类代理及三种带插件记忆模块的Claw类代理,发现当前系统在细粒度关系辨别上依然较弱。我们进一步引入诊断协议,揭示了记忆保存、检索与下游推理阶段的能力差异。
原文摘要 · Abstract (English)
Persistent AI assistants, such as OpenClaw, accumulate large collections of related memories over long-term interactions. As these memories grow, they may reinforce one another, diverge across contexts, or directly conflict, making correct assistance depend on memory relations rather than isolated recall. Existing long-term memory benchmarks rarely probe how agents preserve and utilize such relations during downstream tasks. To address this gap, we introduce SubtleMemory, a benchmark for fine-grained relational memory discrimination in long-running AI agents. SubtleMemory constructs relation-controlled latent semantic artifacts whose variants instantiate complementary, nuanced, or contradictory relations, and embeds them into realistic user-agent histories, requiring agents to recover distributed relational structures during later queries and instructions. The benchmark contains 1,522 evaluation instances over 10 long histories, grounded in 1,090 relation-controlled memory-variant sets and spanning user-related and non-user-related queries. Evaluating six standalone memory systems, two Claw-style agents with native memory modules, and three Claw-style agents with plugin memory modules, we find that current systems remain weak on fine-grained relational memory discrimination. We further introduce diagnostic protocols that reveal distinct capability profiles across memory preservation, retrieval, and downstream reasoning stages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。