首个跨视角视频记忆推理基准,提升智能体空间时间理解能力
EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos

- 设计双视角同步视频记忆推理框架,融合第一人称与第三人称视角
- 提出E²-Select方法,在无训练情况下实现58.2%准确率,超越基线
- 揭示问题提问与答案定位间的视角偏好冲突,凸显跨视角推理挑战
第一人称记忆在具身智能中广泛应用,但可能不足以支撑全面的时空推理。受人类从自身视角和旁观者视角双重回忆的启发,我们提出了EgoExoMem,首个针对同步第一人称与第三人称视频的跨视角记忆推理基准。该基准包含2600个高质量多选题,覆盖八种时序、空间及跨视角问答类型。为支持双视角检索,我们提出E²-Select——一种无需训练的帧选择方法,结合基于相关性的预算分配与每视角k-DPP采样,有效处理视角不对称与跨视角时序一致性问题。实验表明,第一人称与第三人称视角提供互补的记忆线索,而现有多模态大模型仍远未解决此任务:最佳模型仅达55.3%准确率。E²-Select取得58.2%的最先进性能,优于帧选择与RAG-based记忆基线。进一步分析揭示问题表述与答案定位之间存在系统性视角偏好冲突,凸显跨视角记忆推理的新颖性与挑战性。
原文摘要 · Abstract (English)
Egocentric memory is widely used in embodied intelligence, but it may be insufficient for comprehensive spatial-temporal reasoning. Inspired by human recall from both field and observer perspectives, we introduce EgoExoMem, the first benchmark for cross-view memory reasoning over synchronized egocentric and exocentric videos. EgoExoMem contains $2.6K$ high-quality MCQs across eight temporal, spatial, and cross-view QA types. To support dual-view retrieval, we propose E$^2$-Select, a training-free frame selection method for synchronized ego-exo videos. It combines relevance-based budget allocation with per-view k-DPP sampling to handle view asymmetry and cross-view temporal consistency. Experiments show that ego and exo views provide complementary memory cues, while existing MLLMs remain far from solving the benchmark: the best model reaches only $55.3\%$. E$^2$-Select achieves state-of-the-art performance of $58.2\%$ over frame-selection and RAG-based memory baselines. Further analysis reveals systematic view-preference conflicts between question framing and answer grounding, underscoring the novelty and challenge of cross-view memory reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。