arXiv:2501.00358cs.CV2025-01ICCV被引 23

用第一视角视频和传感器构建持久记忆,提升机器人场景理解能力

Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding

  • 结合第一人称视频与深度、位姿等传感器数据构建场景记忆
  • 在多个任务上提升11.7%性能,显著优于现有方法
  • 适合需要长期记忆的机器人交互与操作任务

本文研究从第一人称观测中理解动态3D场景的问题,这是机器人与具身AI的核心挑战。不同于以往仅使用第一人称视频的研究,我们提出基于大语言模型的Embodied VideoAgent,利用第一人称视频和具身传感信息(如深度图与位姿)构建场景记忆。我们进一步设计基于视觉语言模型的方法,在感知到对物体的动作或活动时自动更新记忆。该模型在复杂推理与规划任务中表现优异,在Ego4D-VQ3D上提升4.9%,在OpenEQA上提升5.8%,在EnvQA上提升11.7%。我们还展示了其在生成具身交互和机器人操作感知中的潜力。代码与演示将公开。

原文摘要 · Abstract (English)

This paper investigates the problem of understanding dynamic 3D scenes from egocentric observations, a key challenge in robotics and embodied AI. Unlike prior studies that explored this as long-form video understanding and utilized egocentric video only, we instead propose an LLM-based agent, Embodied VideoAgent, which constructs scene memory from both egocentric video and embodied sensory inputs (e.g. depth and pose sensing). We further introduce a VLM-based approach to automatically update the memory when actions or activities over objects are perceived. Embodied VideoAgent attains significant advantages over counterparts in challenging reasoning and planning tasks in 3D scenes, achieving gains of 4.9% on Ego4D-VQ3D, 5.8% on OpenEQA, and 11.7% on EnvQA. We have also demonstrated its potential in various embodied AI tasks including generating embodied interactions and perception for robot manipulation. The code and demo will be made public.

具身智能场景理解多模态记忆机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。