arXiv:2605.05207cs.CV2026-05被引 4

合成4D动态场景数据集,支持任意像素时空回溯。

Syn4D: A Multiview Synthetic 4D Dataset

论文配图:Syn4D: A Multiview Synthetic 4D Dataset
图 1 · 摘自论文原文
  • 多视角合成动态场景,可精确还原任意时刻任意相机的3D位置。
  • 覆盖4D重建、点跟踪、相机重定位等任务,验证效果优异。
  • 适合研究动态场景理解与时空建模的学者使用。

从单目视频中实现动态场景的稠密3D重建与追踪仍是计算机视觉中的重要开放挑战。该领域的进展受限于高质量、稠密、完整且准确的几何标注数据集稀缺。为此,我们提出Syn4D,一个包含多视角动态场景的合成数据集,提供真实相机运动、深度图、稠密追踪以及参数化人体姿态标注。Syn4D的核心特性是可将任意像素无损地反投影至任意时间与任意相机。我们在多个下游任务中进行了广泛评估,验证了该数据集在4D场景重建、3D点追踪、几何感知相机重定位和人体姿态估计方面的有效性与实用性。实验结果表明,Syn4D具有推动动态场景理解与时空建模研究的巨大潜力。

原文摘要 · Abstract (English)

Dense 3D reconstruction and tracking of dynamic scenes from monocular video remains an important open challenge in computer vision. Progress in this area has been constrained by the scarcity of high-quality datasets with dense, complete, and accurate geometric annotations. To address this limitation, we introduce Syn4D, a multiview synthetic dataset of dynamic scenes that includes ground-truth camera motion, depth maps, dense tracking, and parametric human pose annotations. A key feature of Syn4D is the ability to unproject any pixel into 3D to any time and to any camera. We conduct extensive evaluations across multiple downstream tasks to demonstrate the utility and effectiveness of the proposed dataset, including 4D scene reconstruction, 3D point tracking, geometry-aware camera retargeting, and human pose estimation. The experimental results highlight Syn4D's potential to facilitate research in dynamic scene understanding and spatiotemporal modeling.

4D重建合成数据动态场景姿态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。