让不同步的多视频实现亚帧级同步并重建动态3D场景。
SyncTrack4D: Cross-Video Motion Alignment and Video Synchronization for Multi-Video 4D Gaussian Splatting
- 用4D特征轨迹匹配跨视频运动,通过最优传输对齐时间。
- 在Panoptic Studio上实现0.26帧以下平均时差,PSNR达26.3。
- 无需预设物体或先验模型,适合真实场景多视角动态重建。
由于动态3D场景的高维特性,需融合多视角信息以重建随时间变化的几何与运动。我们提出SyncTrack4D,一种针对真实世界非同步视频集的新型多视频4D Gaussian Splatting(4DGS)方法。该方法直接利用动态场景部分的密集4D轨迹作为线索,同时完成跨视频同步与4DGS重建。首先,通过融合的Gromov-Wasserstein最优传输计算每视频的密集4D特征轨迹及跨视频轨迹对应关系;其次,进行全局帧级时间对齐,最大化匹配4D轨迹的运动重叠;最后,基于运动样条骨架构建多视频4DGS,实现亚帧级同步。最终输出为同步的4DGS表示,包含密集显式3D轨迹与各视频的时间偏移量。我们在Panoptic Studio和SyncNeRF Blender数据集上评估,实现低于0.26帧的平均时差,且在Panoptic Studio上达到26.3 PSNR的高质量4D重建。据我们所知,这是首个无需预设场景对象或先验模型的通用4DGS方法,适用于非同步视频集。
原文摘要 · Abstract (English)
Modeling dynamic 3D scenes is challenging due to their high-dimensional nature, which requires aggregating information from multiple views to reconstruct time-evolving 3D geometry and motion. We present a novel multi-video 4D Gaussian Splatting (4DGS) approach designed to handle real-world, unsynchronized video sets. Our approach, SyncTrack4D, directly leverages dense 4D track representation of dynamic scene parts as cues for simultaneous cross-video synchronization and 4DGS reconstruction. We first compute dense per-video 4D feature tracks and cross-video track correspondences by Fused Gromov-Wasserstein optimal transport approach. Next, we perform global frame-level temporal alignment to maximize overlapping motion of matched 4D tracks. Finally, we achieve sub-frame synchronization through our multi-video 4D Gaussian splatting built upon a motion-spline scaffold representation. The final output is a synchronized 4DGS representation with dense, explicit 3D trajectories, and temporal offsets for each video. We evaluate our approach on the Panoptic Studio and SyncNeRF Blender, demonstrating sub-frame synchronization accuracy with an average temporal error below 0.26 frames, and high-fidelity 4D reconstruction reaching 26.3 PSNR scores on the Panoptic Studio dataset. To the best of our knowledge, our work is the first general 4D Gaussian Splatting approach for unsynchronized video sets, without assuming the existence of predefined scene objects or prior models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。