arXiv:2609.04438cs.CV2026-09

首个专注人物身份长期记忆的多模态评测基准,测试模型对人物关系的跨时间推理能力。

ICM-Bench: Person-Level Identity Reasoning in Multimodal Agents with Long-Term Memory

论文配图:ICM-Bench: Person-Level Identity Reasoning in Multimodal Agents with Long-Term Memory
图 1 · 摘自论文原文
  • 构建包含6位常驻成人、141分钟视频的合成数据集,聚焦人物身份追踪与关联
  • 当前最优模型Gemini 3.1 Pro在需长期身份建模的问题上准确率仅60.3%
  • 揭示现有系统在积累人物相关证据方面仍存在明显不足,适合研究长期记忆的团队使用

长时程多模态智能体不仅需记住事件本身,还需识别参与者身份。这依赖于将反复出现的面容、声音、姓名、人物相关物品、事件及社会关系,持续关联到同一身份。现有长视频和多模态智能体评测侧重广义记忆问答,但未专门评估人物身份的一致性维护与跨时间关系推理。本文提出ICM-Bench(以身份为中心的记忆评测),据我们所知是首个专为评估多模态智能体在长视频记忆中进行身份导向推理而设计的基准。该基准包含839段合成视频(总时长141分钟)和1,217个关于六位常驻成人的开放问题,覆盖一整年生活纪实场景。通过可配置的主题生成管道,自动生成视频并为每道题标注目标身份与可追溯证据。对比了直接图像描述记忆基线、记忆增强型智能体及图检索系统。Gemini 3.1 Pro达到最高总体准确率74.0%,但在需要长期身份画像的问题上下降至60.3%。结果表明,当前系统虽能恢复多数事件记忆,但在围绕稳定人物累积证据时仍不可靠。

原文摘要 · Abstract (English)

Long-horizon multimodal agents should remember not only what happened but also who participated. This capability depends on linking recurring faces, voices, names, person-associated objects, events, and social relations to consistent identities over time. Existing long-video and multimodal-agent benchmarks measure broad memory question answering, but they do not isolate the ability to maintain recurring person identities and reason over their cross-time relations. We introduce ICM-Bench (Identity-Centric Memory Benchmark), which, to the best of our knowledge, is the first benchmark specifically designed to evaluate identity-centric reasoning over long video memories in multimodal agents. The benchmark contains 839 synthetic clips spanning 141 minutes and 1,217 open-ended questions about six recurring adults in a one-year life album. A theme-configurable pipeline generates the video collection and associates each question with its target identities and traceable supporting evidence. We compare direct caption-memory baselines, memory-augmented agents, and graph-retrieval systems. Gemini 3.1 Pro achieves the highest overall accuracy of 74.0%, yet its score falls to 60.3% on questions that require long-term identity profiles. The results show that current systems recover many event-level memories but remain less reliable when evidence must be accumulated around a stable person.

多模态身份推理长时记忆评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。