arXiv:2607.17790cs.CVcs.AI2026-07中稿 · ECCV

仅用单目视频同步重建用户动作与视角变化的4D动态场景。

ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video

论文配图:ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video
图 1 · 摘自论文原文
  • 统一建模视觉、身体、手部与视线运动,无需额外输入。
  • 在多个数据集上实现领先精度与快速推理速度。
  • 适合做沉浸式交互与智能穿戴系统开发的研究者。

可穿戴前向摄像头等第一人称设备为捕捉人类与环境持续交互提供了独特视角。构建一个完整高效的多模态模型来重建这种4D表示极具价值。然而,现有方法常依赖预计算相机轨迹,将场景感知与人体自身运动建模割裂处理,且推理速度慢。为此,我们提出ReViV——首个统一框架,仅通过单目RGB视频即可同时重建观察者与视角动态。该任务被建模为学习包括RGB视频、相机轨迹、凝视方向、全身运动、手部运动和深度在内的多模态信号联合概率分布。基于掩码生成式第一人称变换器,ReViV采用单一前馈架构,实现时空一致的4D重建,具备高速推理能力。在HoloAssist、HOT3D、ARCTIC、Aria Digital Twin及TACO等多个基准上的实验证明,ReViV在整体第一人称身体、手部与凝视重建、相机追踪方面达到最优性能,同时在无重型任务专用先验条件下保持高竞争力的第一人称深度估计。代码与模型已全部开源:https://reviv4d.github.io/。

原文摘要 · Abstract (English)

Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal model capable of reconstructing this 4D representation is therefore highly desirable. However, existing approaches often rely on auxiliary inputs such as pre-computed camera trajectories, treat scene perception and human ego-motion modeling as separate problems despite their strong interdependency, and suffer from slow inference time. To address these limitations, we present ReViV, the first unified framework for holistic egocentric 4D reconstruction that extracts both viewer and view dynamics from a single monocular RGB video. We formulate the task as learning the full joint probability distribution over multimodal signals, including RGB video, camera trajectory, gaze direction, full-body motion, hand motion, and depth. Powered by a Masked Generative Egocentric Transformer, ReViV operates within a single feed-forward architecture to simultaneously reconstruct the temporally consistent 4D reconstruction across the viewer and the view with fast inference speed. Extensive experiments on diverse benchmarks, including HoloAssist, HOT3D, ARCTIC, Aria Digital Twin, and TACO, demonstrate that ReViV achieves state-of-the-art accuracy and efficiency across holistic ego-body, hand, and gaze reconstruction, camera tracking, while maintaining highly competitive egocentric depth estimation without relying on heavy task-specific priors. Code and models are fully open-sourced: https://reviv4d.github.io/.

4D重建第一人称视频多模态建模生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。