评测智能体在多人多模态对话中的记忆能力,发现现有模型表现不佳。
H2HMem: A Multimodal Memory Benchmark for Agents in Human-Human Interactions

- 构建包含多人、多模态的对话记忆评测集
- 多轮实验显示跨参与者和模态记忆严重不足
- 适合研究对话系统、具身智能的开发者参考
大型语言模型代理正越来越多地应用于会议助手、临床记录等人类-人类交互场景,需在对话中观察并保留信息以应对后续查询。与传统人机交互不同,这类环境具有多模态特性,涉及回指、指示等复杂语用现象,且来自多个参与者的异步或矛盾信息交织。然而,现有记忆评测主要聚焦单用户、纯文本场景,无法反映这些挑战。为此,我们提出 H2HMem:一个面向人类-人类交互中多模态记忆能力的基准评测数据集。该数据集包含双人及多方对话,涵盖多模态信息流,并从记忆召回、推理和应用三个维度评估代理性能。对先进代理的实验表明,其在跨模态、跨参与者和跨会话场景下的记忆构建、保持与利用能力存在显著缺陷,揭示了下一代大模型代理的巨大改进空间。
原文摘要 · Abstract (English)
Large language model agents are increasingly deployed in human-human interaction settings, such as meeting assistants and clinical documentation systems, where they must observe conversations and retain information for downstream queries. Unlike traditional human-assistant settings, these environments are inherently multimodal, involve complex discourse phenomena such as anaphora and deixis, and contain asynchronous or conflicting information from multiple participants. However, existing memory benchmarks largely focus on single-user, text-only interactions, failing to capture these challenges. To address this gap, we introduce H2HMem, a Human-to-Human Multimodal Memory Benchmark for evaluating memory capabilities in complex human-human interactions. H2HMem includes both dyadic and multi-party conversations with multimodal information streams, and evaluates agents along three dimensions: memory recall, reasoning, and application. Experiments with advanced agents reveal substantial limitations in constructing, retaining, and utilizing memories across modalities, participants, and sessions, highlighting substantial room for improvement in next-generation LLM agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。