arXiv:2506.06253cs.CV2025-06IJCV综述被引 11

跨视角协同智能:融合第一/第三人称视觉提升机器感知能力

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision

  • 从第一人称与第三人称双视角整合视频理解方法
  • 系统梳理三类协同研究方向及代表性成果
  • 适合关注多视角视频分析与人机认知的读者

从第一人称(内视角)和第三人称(外视角)双重视角感知世界是人类认知的基础,能实现对动态环境的丰富互补理解。近年来,让机器利用这两种视角的协同潜力成为视频理解的重要方向。本文全面综述了内外视角在视频理解中的应用,首先阐述双视角技术集成的实践价值与跨领域合作前景,识别关键研究任务;随后系统归纳为三大研究方向:(1) 利用第一人称数据增强第三人称理解,(2) 利用第三人称数据改进第一人称分析,(3) 融合双视角的联合学习框架。针对每类方向,分析多样任务与相关工作。此外,讨论支持研究的基准数据集,评估其覆盖范围、多样性与适用性。最后,指出当前局限并提出未来研究方向。通过整合双视角洞察,旨在推动视频理解与人工智能发展,使机器更接近人类的感知方式。相关工作汇总见GitHub:https://github.com/ayiyayi/Awesome-Egocentric-and-Exocentric-Vision。

原文摘要 · Abstract (English)

Perceiving the world from both egocentric (first-person) and exocentric (third-person) perspectives is fundamental to human cognition, enabling rich and complementary understanding of dynamic environments. In recent years, allowing the machines to leverage the synergistic potential of these dual perspectives has emerged as a compelling research direction in video understanding. In this survey, we provide a comprehensive review of video understanding from both exocentric and egocentric viewpoints. We begin by highlighting the practical applications of integrating egocentric and exocentric techniques, envisioning their potential collaboration across domains. We then identify key research tasks to realize these applications. Next, we systematically organize and review recent advancements into three main research directions: (1) leveraging egocentric data to enhance exocentric understanding, (2) utilizing exocentric data to improve egocentric analysis, and (3) joint learning frameworks that unify both perspectives. For each direction, we analyze a diverse set of tasks and relevant works. Additionally, we discuss benchmark datasets that support research in both perspectives, evaluating their scope, diversity, and applicability. Finally, we discuss limitations in current works and propose promising future research directions. By synthesizing insights from both perspectives, our goal is to inspire advancements in video understanding and artificial intelligence, bringing machines closer to perceiving the world in a human-like manner. A GitHub repo of related works can be found at https://github.com/ayiyayi/Awesome-Egocentric-and-Exocentric-Vision.

视频理解多视角协同智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。