arXiv:2606.02350cs.CV2026-06

统一重建人、场景与相机,实现全局一致的4D动态感知。

TROPHIES: Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videos

论文配图:TROPHIES: Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videos
图 1 · 摘自论文原文
  • 通过时空推理与人体感知注意力,联合建模动态人与静态场景。
  • 在EgoHuman和EgoExo4D数据集上实现物理合理且全局对齐的4D重建。
  • 适合做多视角动态场景理解与人-环境交互研究的学者使用。

在全局一致的4D空间中重建人类及其周围环境对于全面感知至关重要。然而,以往方法通常依赖单视角输入或解耦人体、场景与相机,难以恢复一致的几何结构、稳定的运动轨迹和物理对齐的路径。为此,我们提出新任务:从多视角视频中统一重建人体、场景与相机,旨在联合估计动态人体、静态场景及相机位姿于同一全局坐标系。我们提出TROPHIES——Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videos,一个专为此任务设计的统一框架。TROPHIES包含人体分支,通过时空推理建模人体;场景分支,利用人体感知注意力重建静态几何。全局对齐与优化模块通过尺度一致性、接触先验与跨视角时间一致性耦合两个分支。在EgoHuman和EgoExo4D数据集上的实验表明,TROPHIES实现了全局对齐、物理合理的4D重建,并在全局保真度与人-场景一致性上持续优于现有范式。

原文摘要 · Abstract (English)

Reconstructing humans and their surrounding environments in a globally consistent 4D space is essential for comprehensive perception. However, prior works typically assume single-view inputs or decouple humans, scenes, and cameras, making them unable to recover coherent geometry, stable motion, and physically aligned trajectories. These limitations motivate us to introduce a new task: unified human-scene-camera reconstruction from multi-view videos, which aims to jointly estimate dynamic humans, static scenes, and camera poses in one global coordinate frame. We propose TROPHIES--Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videos-a unified framework tailored for this task. TROPHIES features a Human Branch that models humans through temporal and spatial reasoning, and a Scene Branch that reconstructs static geometry with human-aware attention. A global alignment and optimization module couples both branches by enforcing scale consistency, contact priors, and cross-view temporal coherence. Experiments on EgoHuman and EgoExo4D demonstrate that TROPHIES achieves globally aligned, physically plausible 4D reconstructions and consistently outperforms existing paradigms in both global fidelity and human-scene consistency.

4D重建多视角视频人体建模场景理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。