arXiv:2607.00157cs.CV2026-07中稿 · ECCV

用单视频高保真重建动物4D动态,跨物种泛化强。

Progressive Pose-Guided 4D Animal Reconstruction from Monocular Video

论文配图:Progressive Pose-Guided 4D Animal Reconstruction from Monocular Video
图 1 · 摘自论文原文
  • 分阶段优化,分离关节动作与非刚性形变
  • 跨物种重建几何精度、时序一致性优于基线
  • 适合动物形态多样、无模板场景的重建任务

从单目视频中重建4D动物极具挑战,源于物种间差异大、关节结构复杂且缺乏可靠模板。现有方法通常依赖特定类别先验以限制泛化能力,或使用无约束生成模型牺牲输入保真度。为此,我们提出基于3D高斯溅射的渐进式测试时优化框架,实现高保真4D动物重建。核心思想是:结合粗略形状先验与渐进策略,可有效分离关节姿态与非刚性形变。具体地,采用对称感知的时间编码,利用双侧线索并缓解相机估计漂移;引入部件条件化的形变机制,由可学习部件锚点和可学习皮肤场引导。大量实验表明,该方法在多种物种间具备稳健泛化能力,在几何精度、时序一致性和视觉保真度上均优于现有基线,即使在先验严重失配情况下依然表现优异。

原文摘要 · Abstract (English)

Reconstructing 4D animals from monocular videos is challenging due to large inter-species variation, complex articulations, and the lack of reliable templates. Existing approaches typically rely on either strict category-specific priors that restrict generalization, or unconstrained generative models that sacrifice input fidelity. To bridge this gap, we present a progressive test-time optimization framework built on 3D Gaussian Splatting for high-fidelity 4D animal reconstruction from a single video. Our key insight is that a coarse shape prior suffices when coupled with a progressive strategy that disentangles articulated pose from non-rigid deformation. Specifically, we employ a symmetry-aware temporal encoding that exploits bilateral cues while absorbing camera estimation drift and a part-conditioned deformation mechanism guided by learnable part anchors and a learnable skinning field. Extensive experiments demonstrate that our approach generalizes robustly across diverse species, achieving superior geometric accuracy, temporal consistency, and visual fidelity compared to existing baselines, even under severe prior mismatch.

4D重建动物建模单目视频高斯溅射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。