用两张无姿态图像实现高精度4D动态重建,无需优化直接输出。
UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images
- 基于动态3D高斯点云的前馈框架,统一估计几何、运动与相机位姿。
- 自监督图像合成损失使多模态联合优化,性能比之前方法提升3倍。
- 适合需要快速高保真4D重建的视觉场景应用,如虚拟拍摄与数字孪生。
从无姿态图像中进行密集4D重建仍是关键挑战,现有方法依赖缓慢的测试时优化或碎片化的专用前馈模型。本文提出UFO-4D,一种统一的前馈框架,仅需一对无姿态图像即可重建稠密显式4D表示。UFO-4D直接估计动态3D高斯点云,实现几何、运动与相机位姿的联合一致估计。核心洞察在于:从单一动态3D高斯表示中可微渲染多种信号,带来显著训练优势。该方法支持自监督图像合成损失,并紧密耦合外观、深度与运动。由于各模态共享相同几何基元,对一者的监督可隐式正则化并提升其他模态。这种协同效应克服数据稀缺问题,使UFO-4D在联合几何、运动与相机位姿估计上相较之前工作最高提升3倍。该表示还支持跨新视角与时间的高保真4D插值。详情见项目页:https://ufo-4d.github.io/
原文摘要 · Abstract (English)
Dense 4D reconstruction from unposed images remains a critical challenge, with current methods relying on slow test-time optimization or fragmented, task-specific feedforward models. We introduce UFO-4D, a unified feedforward framework to reconstruct a dense, explicit 4D representation from just a pair of unposed images. UFO-4D directly estimates dynamic 3D Gaussian Splats, enabling the joint and consistent estimation of 3D geometry, 3D motion, and camera pose in a feedforward manner. Our core insight is that differentiably rendering multiple signals from a single Dynamic 3D Gaussian representation offers major training advantages. This approach enables a self-supervised image synthesis loss while tightly coupling appearance, depth, and motion. Since all modalities share the same geometric primitives, supervising one inherently regularizes and improves the others. This synergy overcomes data scarcity, allowing UFO-4D to outperform prior work by up to 3 times in joint geometry, motion, and camera pose estimation. Our representation also enables high-fidelity 4D interpolation across novel views and time. Please visit our project page for visual results: https://ufo-4d.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。