单目视频实时重建4D场景,几何与运动联合建模
4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere
- 编码一次,任意时间任意帧可查询3D几何和运动
- 在多个4D重建任务上超越现有方法
- 适合需要动态场景全时序重建的研究者
我们提出4RC,一种统一的前馈框架,用于从单目视频中进行4D重建。与以往将运动与几何解耦或仅生成稀疏轨迹、双视图光流等有限4D属性的方法不同,4RC学习一个整体的4D表示,联合捕捉密集场景几何和运动动态。其核心是引入一种新型的‘编码一次,任意时间任意位置查询’范式:通过Transformer主干将整个视频编码为紧凑的时空潜在空间,条件解码器可高效查询任意目标时间戳下的任意帧的3D几何与运动。为促进学习,我们以最小化因子分解形式表示每视角4D属性,将其分解为基础几何与随时间变化的相对运动。大量实验表明,4RC在多种4D重建任务中均优于先前及同期方法。
原文摘要 · Abstract (English)
We present 4RC, a unified feed-forward framework for 4D reconstruction from monocular videos. Unlike existing approaches that typically decouple motion from geometry or produce limited 4D attributes such as sparse trajectories or two-view scene flow, 4RC learns a holistic 4D representation that jointly captures dense scene geometry and motion dynamics. At its core, 4RC introduces a novel encode-once, query-anywhere and anytime paradigm: a transformer backbone encodes the entire video into a compact spatio-temporal latent space, from which a conditional decoder can efficiently query 3D geometry and motion for any query frame at any target timestamp. To facilitate learning, we represent per-view 4D attributes in a minimally factorized form by decomposing them into base geometry and time-dependent relative motion. Extensive experiments demonstrate that 4RC outperforms prior and concurrent methods across a wide range of 4D reconstruction tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。