无需相机位姿即可实现多视角动态场景的实时重建。
No Pose, No Problem in 4D: Feed-Forward Dynamic Gaussians from Unposed Multi-View Videos

- 将高斯点运动分解为图像平面位移与深度变化,直接用伪光流监督。
- 在4个动态数据集上超越已有前馈方法,速度比逐场景优化快数个量级。
- 适合需要快速重建动态场景的实时应用,如AR/VR和机器人导航。
现有前馈式3D高斯溅射方法在单方面重建上进展显著,但尚无方法能在一次前馈中同时处理动态内容、多视角输入和未知相机位姿。现有方法要么需精确相机位姿,要么仅支持单目输入;无位姿的多视角方法仅适用于静态场景;而逐场景优化虽弥补部分缺陷,但每场景耗时数分钟至数小时。本文提出NoPo4D,首个解决该空白的前馈系统。基于预训练几何主干与近期4D高斯框架,引入速度分解机制,将高斯运动拆分为像素级图像平面偏移与深度变化,从而直接利用伪真值光流监督2D分量。此设计规避了依赖可微渲染的位姿精度要求,也免去对3D运动真值的依赖。系统还配备双向运动编码器实现跨视角与跨帧特征聚合,以及视图相关不透明度以缓解高斯点在时空上的错位问题。在四个多视角动态基准上,NoPo4D持续优于先前前馈基线,经可选后优化阶段更超越逐场景优化方法,且运行速度快数个量级。
原文摘要 · Abstract (English)
Recent feed-forward 3D gaussian splatting methods have made dramatic progress on individual aspects of 3D scene reconstruction, but no existing method jointly addresses dynamic content, multi-view input, and unknown camera poses in a single feed-forward pass. Methods that handle dynamics either require accurate camera poses or accept only monocular input; pose-free multi-view methods address only static scenes; and per-scene optimization methods bridge some of these gaps but at minutes-to-hours cost per scene. We introduce NoPo4D, the first feed-forward system that addresses this empty quadrant. Building on a pretrained geometry backbone and recent 4D Gaussian frameworks, NoPo4D introduces a velocity decomposition that splits Gaussian motion into per-pixel image-plane shifts and depth changes, allowing direct supervision from pseudo ground-truth optical flow on the 2D component. This sidesteps both the differentiable rendering that couples prior posed methods to pose accuracy and the 3D motion ground truth that prior pose-free methods require. The system is rounded out by a bidirectional motion encoder for cross-view and cross-frame feature aggregation, and view-dependent opacity that mitigates cross-view and cross-timestep Gaussian misalignments. On four multi-view dynamic benchmarks, NoPo4D consistently outperforms prior feed-forward baselines, and with an optional post-optimization stage surpasses per-scene optimization methods, while running orders of magnitude faster.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。