测试多源异构信息如何被智能体整合推理。
SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory

- 设计跨源多模态证据检索与对齐机制
- 1877个样本中现有系统表现不佳
- 适合研究多模态记忆与智能体推理的学者
现有多模态记忆推理基准大多在预组装上下文中评估系统,但忽略了智能体从独立来源分散证据的能力。我们指出,跨源多模态记忆组合是多模态智能体中的重要且未被充分考察的瓶颈,尤其当相关证据分散在对话、个人资料、截图、表格、图像和文档等异构载体中时。为此,我们提出源分布多模态记忆基准(SMMBench),用于评估智能体能否从多个来源中检索、对齐并组合分散的多模态证据,而非仅在单一整理好的上下文中推理。该基准涵盖四项核心能力:(1) 跨源多模态推理;(2) 冲突解决;(3) 偏好推理;(4) 基于记忆的动作预测。数据集包含1877个样本,源自264个不同来源。在代表性记忆型与基于检索的基线模型上实验表明,当前系统在这些能力上仍表现困难,凸显源分布多模态记忆作为关键挑战的重要性。数据已公开于 https://huggingface.co/datasets/HuacanChai/SMMBench。
原文摘要 · Abstract (English)
Existing benchmarks for multimodal memory reasoning largely evaluate systems within pre-assembled contexts, but under-evaluate whether agents can use evidence distributed across independently originated sources. We argue that source-distributed memory composition is an important and under-examined bottleneck in multimodal agent memory, especially when relevant evidence is fragmented across heterogeneous artifacts such as conversations, profiles, screenshots, tables, images, and documents. To address this gap, we introduce Source-distributed Multimodal Memory Benchmark(SMMBench), which measures whether agents can retrieve, align, and compose multimodal evidence scattered across multiple sources rather than reason within a single curated context. SMMBench evaluates four core capabilities: (1) cross-source multimodal reasoning; (2) conflict resolution; (3) preference reasoning; (4) memory-grounded action prediction. The benchmark contains 1877 samples grounded in 264 sources. Experiments on representative memory-style and retrieval-based baselines show that current systems still struggle on these capabilities, positioning source-distributed multimodal memory as an important and still under-evaluated challenge for multimodal agents. Our data are available at https://huggingface.co/datasets/HuacanChai/SMMBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。