测试AI记忆系统在多人社交场景下的表现,发现现有模型普遍失效。
SocialMemBench: Are AI Memory Systems Ready for Social Group Settings?

- 构建五类社交网络的合成数据集,模拟真实群体互动记忆需求。
- 主流AI记忆框架在多人群体中准确率仅0.12-0.18,远低于人类水平。
- 适用于社交助理、群组协作等需理解社会关系的智能应用。
当前AI助手的记忆系统专为单人对话设计,在多人社交场景中表现失常。这一差距对今日社交助理与主动个人助手至关重要——后者需整合用户的社交背景。现有基准仅覆盖两人或职场对话,未针对多人群体设计,而后者要求记忆锚定共享历史,区分群体规范与个体例外,并在成员退出后仍能正确归属信息。我们提出SocialMemBench,包含五种社交原型(亲密朋友、家庭、休闲、兴趣社群、熟人圈)和三类规模(4-30人),共430个角色、7,355轮对话,生成1,031个问答对,涵盖九类问题以检验不同架构能力。五个已知失效模式(单流混淆、时间状态覆盖、大规模实体合并、跨角色知识缺失、规范与个体混淆)均可验证;两项研究探针提供证据,三项仍待验证。全上下文Gemini 2.5 Flash参考模型在小网络上仅达0.721,低于盲评推理模型平均0.98,表明该基准极具挑战性,即使完全访问对话内容亦难应对。在全部43个网络中,四个开源记忆框架(Mem0、LangMem、Graphiti、Cognee)得分集中在0.12-0.18之间,95%置信区间重叠,显著低于无压缩检索参考值0.345及匹配回答者全上下文参考值0.369(GPT-4o-mini)。当前记忆系统存在明显能力鸿沟。
原文摘要 · Abstract (English)
Memory systems for AI assistants were built for single-user dialogue and fail characteristically when applied to multi-party social group settings. This gap matters for the social assistants being built today: group-acting agents embedded in chat platforms, and proactive personal-assistant agents whose holistic model of a user must include their social context. Existing memory benchmarks evaluate dyadic or workplace dialogue; none targets multi-party social groups, where memory must anchor facts in shared history rather than professional roles, separate group norms from individual exceptions, and correctly attribute even after member departure. We introduce SocialMemBench, a benchmark of human-verified synthetic social group networks across five archetypes (close friends, family, recreational, interest community, acquaintance network) and three group-size tiers (4-30 members), with 430 personas and 7,355 conversation turns, yielding 1,031 QA pairs across nine question categories. Each category isolates an architectural capability, and the five failure modes (single-stream conflation, temporal-state overwrite, entity merging at scale, missing cross-persona knowledge, norm-individual conflation) are testable hypotheses; our two research probes Subject-Mem and SMG provide evidence on two, three remain open. A full-context Gemini 2.5 Flash reference reaches only 0.721 against a blind-critic reasoning-model mean of 0.98 on small networks, indicating the benchmark is genuinely difficult even with complete access to the conversation. Across all 43 networks, the four open-source memory frameworks evaluated (Mem0, LangMem, Graphiti, Cognee) cluster in the 0.12-0.18 question-weighted range with overlapping 95% CIs, well below an uncompressed retrieval reference of 0.345 and a matched-answerer full-context reference of 0.369 (GPT-4o-mini). Current memory systems show a measurable gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。