arXiv:2506.17896cs.CVcs.AI2025-06中稿 · ICLR被引 8

用多视角数据生成第一人称视角图像,提升机器人与元宇宙应用效果

EgoWorld: Translating Exocentric View to Egocentric View using Rich Exocentric Observations

  • 融合点云、3D手姿和文本,从第三人称视角重建第一人称视图
  • 在4个数据集上达到顶尖性能,对新物体和动作泛化能力强
  • 适用于真实场景,适合做机器人操作和虚拟现实系统开发

第一人称视觉对人类和机器的视觉理解至关重要,尤其在捕捉操作任务中的手物交互细节方面。将第三人称视角转换为第一人称视角可显著提升增强现实(AR)、虚拟现实(VR)和机器人应用效果。然而,现有方法受限于依赖2D线索、需同步多视角设置,以及推理时需初始第一人称帧和相对相机位姿等不切实际假设。为此,我们提出EgoWorld,一种新颖框架,通过丰富的第三人称观测(包括点云、3D手姿态和文本描述)重建第一人称视角。该方法先从估计的第三人称深度图重建点云,再重投影至第一人称视角,最后利用扩散模型生成密集且语义一致的第一人称图像。在四个数据集(H2O、TACO、Assembly101、Ego-Exo4D)上的评估表明,EgoWorld表现领先,对新物体、动作、场景和主体具有强泛化能力。此外,其在真实场景样本上也表现出鲁棒性,凸显实际应用价值。

原文摘要 · Abstract (English)

Egocentric vision is essential for both human and machine visual understanding, particularly in capturing the detailed hand-object interactions needed for manipulation tasks. Translating third-person views into first-person views significantly benefits augmented reality (AR), virtual reality (VR) and robotics applications. However, current exocentric-to-egocentric translation methods are limited by their dependence on 2D cues, synchronized multi-view settings, and unrealistic assumptions such as the necessity of an initial egocentric frame and relative camera poses during inference. To overcome these challenges, we introduce EgoWorld, a novel framework that reconstructs an egocentric view from rich exocentric observations, including point clouds, 3D hand poses, and textual descriptions. Our approach reconstructs a point cloud from estimated exocentric depth maps, reprojects it into the egocentric perspective, and then applies diffusion model to produce dense, semantically coherent egocentric images. Evaluated on four datasets (i.e., H2O, TACO, Assembly101, and Ego-Exo4D), EgoWorld achieves state-of-the-art performance and demonstrates robust generalization to new objects, actions, scenes, and subjects. Moreover, EgoWorld exhibits robustness on in-the-wild examples, underscoring its practical applicability. Project page is available at https://redorangeyellowy.github.io/EgoWorld/.

第一人称视觉视角转换扩散模型机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。