用几何约束提升单目视频动态场景重建精度与一致性
SirenPose: Dynamic Scene Reconstruction via Geometric Supervision
- 结合正弦网络周期性与关键点几何监督,增强时空一致性
- 在DAVIS上降低17.8% FVD、28.7% FID,LPIPS提升6.0%
- 适合需要高精度动态重建与运动平滑的视觉应用
我们提出SirenPose,一种融合正弦表示网络周期性激活特性与基于关键点的几何监督的几何感知损失函数,可从单目视频中实现精确且时间一致的动态3D场景重建。现有方法在快速运动、多物体交互、遮挡和快速场景变化等挑战性场景下常出现运动保真度低、时空不连贯的问题。SirenPose引入物理启发的约束,确保关键点预测在空间和时间维度上的协同性,同时利用高频信号建模捕捉精细几何细节。我们还将UniKPT数据集扩展至60万标注实例,并引入图神经网络建模关键点关系与结构关联。在Sintel、Bonn和DAVIS等基准测试中,SirenPose持续优于当前最优方法。在DAVIS上,相比MoSCA,FVD降低17.8%,FID降低28.7%,LPIPS提升6.0%;同时显著提升时间一致性、几何精度、用户评分与运动平滑性。在姿态估计方面,相比Monst3R,SirenPose在绝对轨迹误差、平移与旋转相对姿态误差上均更低,凸显其在处理快速运动、复杂动态与物理合理重建方面的有效性。
原文摘要 · Abstract (English)
We introduce SirenPose, a geometry-aware loss formulation that integrates the periodic activation properties of sinusoidal representation networks with keypoint-based geometric supervision, enabling accurate and temporally consistent reconstruction of dynamic 3D scenes from monocular videos. Existing approaches often struggle with motion fidelity and spatiotemporal coherence in challenging settings involving fast motion, multi-object interaction, occlusion, and rapid scene changes. SirenPose incorporates physics inspired constraints to enforce coherent keypoint predictions across both spatial and temporal dimensions, while leveraging high frequency signal modeling to capture fine grained geometric details. We further expand the UniKPT dataset to 600,000 annotated instances and integrate graph neural networks to model keypoint relationships and structural correlations. Extensive experiments on benchmarks including Sintel, Bonn, and DAVIS demonstrate that SirenPose consistently outperforms state-of-the-art methods. On DAVIS, SirenPose achieves a 17.8 percent reduction in FVD, a 28.7 percent reduction in FID, and a 6.0 percent improvement in LPIPS compared to MoSCA. It also improves temporal consistency, geometric accuracy, user score, and motion smoothness. In pose estimation, SirenPose outperforms Monst3R with lower absolute trajectory error as well as reduced translational and rotational relative pose error, highlighting its effectiveness in handling rapid motion, complex dynamics, and physically plausible reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。