arXiv:2410.20621cs.CV2024-10中稿 · Computer Vision an…综述被引 16

对比第一人称与第三人称视觉,梳理联合建模新方向

Egocentric and Exocentric Methods: A Short Survey

  • 融合第一人称与第三人称视角数据进行联合建模
  • 揭示双视角互补信号对视频理解的提升作用
  • 适合关注多视角视频分析的科研人员参考

第一人称视觉捕捉佩戴者视角的场景,第三人称视觉则提供全局场景上下文。联合建模第一人称与第三人称视角对于发展下一代AI智能体至关重要。近年来,社区重新关注第一人称视觉领域。尽管第三人称和第一人称视角各自已有深入研究,但同步研究两者的工作仍很少。第三人称视频包含可迁移至第一人称视频的丰富信息。本文系统综述了第一人称与第三人称视觉融合的研究进展,涵盖关键数据集与核心应用场景,梳理了最新技术突破。通过呈现当前研究现状,我们认为该短篇综述对视频理解领域具有重要参考价值,尤其在多视角建模至关重要的场景中。

原文摘要 · Abstract (English)

Egocentric vision captures the scene from the point of view of the camera wearer, while exocentric vision captures the overall scene context. Jointly modeling ego and exo views is crucial to developing next-generation AI agents. The community has regained interest in the field of egocentric vision. While the third-person view and first-person have been thoroughly investigated, very few works aim to study both synchronously. Exocentric videos contain many relevant signals that are transferrable to egocentric videos. This paper provides a timely overview of works combining egocentric and exocentric visions, a very new but promising research topic. We describe in detail the datasets and present a survey of the key applications of ego-exo joint learning, where we identify the most recent advances. With the presentation of the current status of the progress, we believe this short but timely survey will be valuable to the broad video-understanding community, particularly when multi-view modeling is critical.

多视角视频理解视觉融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。