arXiv:2503.08344cs.CV2025-03CVPR

将第一人称视频分解为动态与静态成分,提升环境理解能力。

DIV-FF: Dynamic Image-Video Feature Fields For Environment Understanding in Egocentric Videos

  • 分离场景中的持久、动态与动作主体成分,融合图像与视频语言特征。
  • 在动态变化场景中显著优于现有方法,实现时序一致的环境理解。
  • 适合长期时空场景理解任务,如机器人导航与增强现实应用。

第一人称视频中的环境理解对机器人、增强现实和辅助技术等应用至关重要。这类视频具有动态交互性强、与佩戴者行为紧密相关的特点。传统方法通常仅处理孤立片段,或无法有效整合丰富的语义与几何信息,限制了场景认知能力。本文提出动态图像-视频特征场(DIV-FF)框架,将第一人称场景分解为持久、动态及基于动作主体的组件,并融合图像与视频语言特征。该模型支持细粒度分割,捕捉物体可操作性,理解周围环境,并保持时间上的一致性。DIV-FF在动态演化场景中表现优于当前最优方法,展现出推动长期时空场景理解的潜力。

原文摘要 · Abstract (English)

Environment understanding in egocentric videos is an important step for applications like robotics, augmented reality and assistive technologies. These videos are characterized by dynamic interactions and a strong dependence on the wearer engagement with the environment. Traditional approaches often focus on isolated clips or fail to integrate rich semantic and geometric information, limiting scene comprehension. We introduce Dynamic Image-Video Feature Fields (DIV FF), a framework that decomposes the egocentric scene into persistent, dynamic, and actor based components while integrating both image and video language features. Our model enables detailed segmentation, captures affordances, understands the surroundings and maintains consistent understanding over time. DIV-FF outperforms state-of-the-art methods, particularly in dynamically evolving scenarios, demonstrating its potential to advance long term, spatio temporal scene understanding.

第一人称视频场景理解动态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。