arXiv:2507.10437cs.CV2025-07被引 8

无需关键点标注,从视频重建可动画3D动物

4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos

  • 用密集特征网络直接从2D图像映射到SMAL参数
  • 融合轮廓、部件、像素和时序线索,重建更准确连贯
  • 适合需要批量生成高质量3D动物资产的研究者

现有从视频重建可动画3D动物的方法通常依赖稀疏语义关键点来拟合参数化模型,但获取这些关键点耗时费力,且在有限动物数据上训练的关键点检测器往往不可靠。为此,我们提出4D-Animal,一种无需稀疏关键点标注即可从视频重建可动画3D动物的新框架。该方法引入一个密集特征网络,将2D表示映射到SMAL参数,提升了拟合过程的效率与稳定性。此外,我们设计了一种分层对齐策略,整合预训练2D视觉模型提供的轮廓、部件级、像素级和时序线索,实现跨帧精确且时间一致的重建。大量实验表明,4D-Animal优于基于模型和无模型的基线方法。此外,本方法生成的高质量3D资产可服务于其他3D任务,展现出大规模应用潜力。代码已公开于https://github.com/zhongshsh/4D-Animal。

原文摘要 · Abstract (English)

Existing methods for reconstructing animatable 3D animals from videos typically rely on sparse semantic keypoints to fit parametric models. However, obtaining such keypoints is labor-intensive, and keypoint detectors trained on limited animal data are often unreliable. To address this, we propose 4D-Animal, a novel framework that reconstructs animatable 3D animals from videos without requiring sparse keypoint annotations. Our approach introduces a dense feature network that maps 2D representations to SMAL parameters, enhancing both the efficiency and stability of the fitting process. Furthermore, we develop a hierarchical alignment strategy that integrates silhouette, part-level, pixel-level, and temporal cues from pre-trained 2D visual models to produce accurate and temporally coherent reconstructions across frames. Extensive experiments demonstrate that 4D-Animal outperforms both model-based and model-free baselines. Moreover, the high-quality 3D assets generated by our method can benefit other 3D tasks, underscoring its potential for large-scale applications. The code is released at https://github.com/zhongshsh/4D-Animal.

3D重建动物建模视频驱动SMAL

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。