arXiv:2512.06158cs.CV2025-12被引 1

用追踪引导生成4D动态模型,解决跨视角与时间的连贯性问题。

Tracking-Guided 4D Generation: Foundation-Tracker Motion Priors for 3D Model Animation

  • 通过融合追踪器运动先验,增强扩散模型中间特征的时间一致性
  • 在多视图视频与4D重建任务中均超越基线,生成稳定可编辑的4D资产
  • 适用于需要高保真动态3D内容的动画、游戏与数字孪生领域

从稀疏输入生成动态4D物体极具挑战,因需在多视角与时间维度上同时保持外观与运动的一致性,并抑制伪影与时序漂移。我们假设,视角差异源于仅依赖像素或潜在空间视频扩散损失的监督,缺乏显式的时序感知、特征级追踪引导。本文提出两阶段框架Track4DGen,将多视图视频扩散模型与基础点追踪器及混合4D高斯溅射(4D-GS)重建器耦合。核心思想是将追踪器生成的运动先验注入扩散生成与4D-GS的中间特征表示中。第一阶段,在扩散生成器内强制实现密集特征级点对应,生成时序一致的特征,有效抑制外观漂移并提升跨视角一致性。第二阶段,采用混合运动编码重建4D-GS,将共位扩散特征(携带第一阶段追踪先验)与Hex-plane特征拼接,并引入4D球谐函数以提升动态建模精度。Track4DGen在多视图视频生成与4D生成基准测试中均超越基线,生成时序稳定、支持文本编辑的4D资产。最后,我们构建了高质量数据集Sketchfab28,用于物体中心的4D生成评估与未来研究。

原文摘要 · Abstract (English)

Generating dynamic 4D objects from sparse inputs is difficult because it demands joint preservation of appearance and motion coherence across views and time while suppressing artifacts and temporal drift. We hypothesize that the view discrepancy arises from supervision limited to pixel- or latent-space video-diffusion losses, which lack explicitly temporally aware, feature-level tracking guidance. We present \emph{Track4DGen}, a two-stage framework that couples a multi-view video diffusion model with a foundation point tracker and a hybrid 4D Gaussian Splatting (4D-GS) reconstructor. The central idea is to explicitly inject tracker-derived motion priors into intermediate feature representations for both multi-view video generation and 4D-GS. In Stage One, we enforce dense, feature-level point correspondences inside the diffusion generator, producing temporally consistent features that curb appearance drift and enhance cross-view coherence. In Stage Two, we reconstruct a dynamic 4D-GS using a hybrid motion encoding that concatenates co-located diffusion features (carrying Stage-One tracking priors) with Hex-plane features, and augment them with 4D Spherical Harmonics for higher-fidelity dynamics modeling. \emph{Track4DGen} surpasses baselines on both multi-view video generation and 4D generation benchmarks, yielding temporally stable, text-editable 4D assets. Lastly, we curate \emph{Sketchfab28}, a high-quality dataset for benchmarking object-centric 4D generation and fostering future research.

4D生成运动追踪扩散模型高斯溅射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。