arXiv:2606.00825cs.CVcs.ET2026-06被引 2

构建首个长时记忆视觉问答数据集,测试AI眼镜在真实生活中的记忆能力。

SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory

论文配图:SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory
图 1 · 摘自论文原文
  • 基于52.9小时第一视角视频,构建多模态长时记忆问答对
  • 4,853个问题覆盖物体、位置、意图等六大记忆类型,含不可回答选项防幻觉
  • 验证主流模型仍难胜任真实场景记忆任务,适合研究长期记忆与具身智能者

AI眼镜为个性化记忆助手提供了理想平台。为真正实用,系统需超越短期视频理解,解决人类在长期第一视角视频中因实际、个人或社交需求产生的记忆空白。现有第一视角数据集多聚焦动作识别或短片段通用问答,仅衡量感知能力而非真实记忆需求。我们提出SuperMemory-VQA,一个面向长时记忆任务的视觉问答基准。该数据集包含52.9小时日常活动记录,同步提供RGB视频、语音转录、眼动、IMU及SLAM轨迹。通过人工验证标注流程,构建了4,853个基于真实场景的问答对,涵盖物体与位置记忆、意图回溯、视觉场景重建、时间线复原、对话记忆及上下文检索。每个问题为多选题,并设明确“不可回答”选项以检验幻觉鲁棒性。对主流代理框架与大模型的基准测试表明,当前系统在真实记忆任务上仍不可靠,凸显了需发展能仅在证据充分时作答的具身记忆架构。参与者调查显示,问题具有现实性、实用性与日常记忆契合度。

原文摘要 · Abstract (English)

AI glasses present a compelling platform for AI agents to serve as personalized memory assistants. To be genuinely useful, such systems must move beyond short-term video comprehension and address memory gaps that humans experience for practical, personal, or social purposes over longitudinal egocentric video streams. However, existing egocentric datasets predominantly focus on action recognition or generic QAs from short clips, measuring perceptual capabilities rather than realistic human memory needs. We introduce SuperMemory-VQA, an egocentric visual question answering (VQA) dataset for evaluating AI assistants on practical, long-horizon memory tasks. It contains 52.9 hours of everyday activities recorded with AI glasses, including synchronized RGB video, audio transcription, eye gaze, IMU, and SLAM trajectories. Through a human-verified annotation pipeline, we construct grounded 4,853 question-answer pairs that span object and location memory, intent recall, visual scene recall, timeline reconstruction, conversational memory, and in-context retrieval. Each question is posed as multiple-choice with an explicit "unanswerable" option to test hallucination robustness. Benchmarking leading agentic frameworks and LLM backbones reveals that existing systems remain far from reliable on real-world memory tasks, highlighting the need for new architectures for grounded AI memory that can answer only when evidence is sufficient. A participant survey further supports that our questions are realistic, useful, and aligned with everyday memory needs.

视觉问答长时记忆第一人称视频AI眼镜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。