arXiv:2509.18436cs.AIcs.CL2025-09EMNLP被引 8

让AI从多模态记忆中回答视觉回忆问题,提升真实场景下的记忆检索能力。

Memory-QA: Answering Recall Questions Based on Multimodal Memories

论文配图:Memory-QA: Answering Recall Questions Based on Multimodal Memories
图 1 · 摘自论文原文
  • 构建专用记忆增强与时空感知的多信号检索流程
  • 在多记忆问答任务上实现最高14%的准确率提升
  • 适合需要长期记忆与环境上下文理解的应用场景

我们提出Memory-QA这一新型现实任务,要求基于先前存储的多模态记忆回答关于视觉内容的回忆问题。该任务面临独特挑战:需生成任务导向的记忆、有效利用记忆中的时间与位置信息,并能综合多个记忆回答问题。为此,我们设计了Pensieve全流程框架,整合记忆特异性增强、时序与位置感知的多信号检索以及多记忆QA微调。我们建立了一个多模态基准数据集,展示该任务中的各类现实挑战,并证明Pensieve在多项指标上显著优于现有方案,最高可提升14%的问答准确率。

原文摘要 · Abstract (English)

We introduce Memory-QA, a novel real-world task that involves answering recall questions about visual content from previously stored multimodal memories. This task poses unique challenges, including the creation of task-oriented memories, the effective utilization of temporal and location information within memories, and the ability to draw upon multiple memories to answer a recall question. To address these challenges, we propose a comprehensive pipeline, Pensieve, integrating memory-specific augmentation, time- and location-aware multi-signal retrieval, and multi-memory QA fine-tuning. We created a multimodal benchmark to illustrate various real challenges in this task, and show the superior performance of Pensieve over state-of-the-art solutions (up to 14% on QA accuracy).

多模态记忆问答系统时空感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。