让可穿戴设备实时记忆并定位看到的物体,提升智能眼镜实用性。
Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory
- 在线处理视频流,用紧凑记忆存储关键信息,不依赖完整视频。
- 在Ego4D数据集上最高定位成功率仅约4%,但全知条件下可达81.92%。
- 适合研究实时视觉记忆、可穿戴设备与高效目标追踪的开发者。
情景记忆检索使可穿戴摄像头能够回忆先前视频中观察到的物体或事件。然而,现有方法假设查询时可访问完整视频(离线设置),限制了其在功耗和存储受限的可穿戴设备上的实际应用。为构建更实用的情景记忆系统,我们提出在线视觉查询二维任务(OVQ2D):模型需在线处理视频流,每帧仅观察一次,并利用紧凑记忆而非完整历史进行物体定位。为此,我们设计了面向自指视频流的对象记忆框架ESOM,集成对象发现、跟踪与记忆模块,用于高效存储时空对象信息。在Ego4D数据集上的实验表明,相比其他在线方法,ESOM表现更优,但整体任务仍具挑战性,最佳性能约为4%。当对象跟踪或发现完全准确时,精度分别提升至31.91%和40.55%,两者均完美时达81.92%,凸显对相关组件深入研究的必要性。
原文摘要 · Abstract (English)
Episodic memory retrieval enables wearable cameras to recall objects or events previously observed in video. However, existing formulations assume an "offline" setting with full video access at query time, limiting their applicability in real-world scenarios with power and storage-constrained wearable devices. Towards more application-ready episodic memory systems, we introduce Online Visual Query 2D (OVQ2D), a task where models process video streams online, observing each frame only once, and retrieve object localizations using a compact memory instead of full video history. We address OVQ2D with ESOM (Egocentric Streaming Object Memory), a novel framework integrating an object discovery module, an object tracking module, and a memory module that find, track, and store spatio-temporal object information for efficient querying. Experiments on Ego4D demonstrate ESOM's superiority over other online approaches, though OVQ2D remains challenging, with top performance at only ~4% success. ESOM's accuracy increases markedly with perfect object tracking (31.91%), discovery (40.55%), or both (81.92%), underscoring the need of applied research on these components.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。