无需扩散模型,单视频实现360°动态物体精确重建
4DGS360: 360° Gaussian Reconstruction of Dynamic Objects from a Single Video
- 用3D原生初始化解决遮挡区域几何模糊问题
- 在iPhone360等数据集上达到当前最佳性能
- 适合做动态场景三维重建的研究者和开发者
我们提出4DGS360,一种无需扩散模型的框架,可从单个视角视频中实现360°动态物体的重建。现有方法因过度依赖2D先验,在训练视图中易对可见表面过拟合,导致360°几何不一致。4DGS360通过先进的3D原生初始化缓解遮挡区域的几何歧义。其提出的3D跟踪器AnchorTAP3D,利用可靠的2D跟踪点作为锚点,强化3D点轨迹,抑制漂移,并提供可靠初始状态,从而保留遮挡区域的几何结构。该初始化与优化结合后,生成一致的360° 4D重建结果。我们还构建了新基准iPhone360,测试相机与训练视图最大夹角达135°,支持现有数据集无法提供的360°评估。实验表明,4DGS360在iPhone360、iPhone和DAVIS数据集上均达到最优表现,定性与定量结果俱佳。
原文摘要 · Abstract (English)
We introduce 4DGS360, a diffusion-free framework for 360$^{\circ}$ dynamic object reconstruction from casual monocular video. Existing methods often fail to reconstruct consistent 360$^{\circ}$ geometry, as their heavy reliance on 2D-native priors causes initial points to overfit to visible surface in each training view. 4DGS360 addresses this challenge through a advanced 3D-native initialization that mitigates the geometric ambiguity of occluded regions. Our proposed 3D tracker, AnchorTAP3D, produces reinforced 3D point trajectories by leveraging confident 2D track points as anchors, suppressing drift and providing reliable initialization that preserves geometry in occluded regions. This initialization, combined with optimization, yields coherent 360$^{\circ}$ 4D reconstructions. We further present iPhone360, a new benchmark where test cameras are placed up to 135$^{\circ}$ apart from training views, enabling 360$^{\circ}$ evaluation that existing datasets cannot provide. Experiments show that 4DGS360 achieves state-of-the-art performance on the iPhone360, iPhone, and DAVIS datasets, both qualitatively and quantitatively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。