将马匹4D重建拆解为运动与外观两部分,提升重建精度与效率。
4DEquine: Disentangling Motion and Appearance for 4D Equine Reconstruction from Monocular Video
- 分离运动与外观重建,用时空变换器和前馈网络分别处理
- 仅需单张图像即可生成高保真可动画3D高斯化身,真实数据上达顶尖水平
- 自建合成数据集支持训练,适合动物行为分析与虚拟仿真研究
从单目视频进行马匹类动物的4D重建对动物福利研究具有重要意义。现有主流方法需联合优化整段视频的运动与外观,计算耗时且易受观测不全影响。本文提出4DEquine框架,将4D重建问题分解为动态运动重建与静态外观重建两个子问题。针对运动,设计一种简单有效的时空变换器,并引入后优化阶段,从视频中回归出平滑且像素对齐的姿态与形状序列;针对外观,构建新型前馈网络,仅需单张图像即可重建出高保真、可动画化的3D Gaussian Avatar。为辅助训练,构建大规模合成运动数据集VarenPoser(包含高质量表面运动与多样相机轨迹)及合成外观数据集VarenTex(通过多视角扩散模型生成的真实感多视图图像)。仅在合成数据上训练,4DEquine在真实世界APT36K与AiM数据集上均达到当前最优性能,验证了该框架及新数据集在几何与外观重建上的优越性。全面消融实验进一步证明了运动与外观重建网络的有效性。
原文摘要 · Abstract (English)
4D reconstruction of equine family (e.g. horses) from monocular video is important for animal welfare. Previous mainstream 4D animal reconstruction methods require joint optimization of motion and appearance over a whole video, which is time-consuming and sensitive to incomplete observation. In this work, we propose a novel framework called 4DEquine by disentangling the 4D reconstruction problem into two sub-problems: dynamic motion reconstruction and static appearance reconstruction. For motion, we introduce a simple yet effective spatio-temporal transformer with a post-optimization stage to regress smooth and pixel-aligned pose and shape sequences from video. For appearance, we design a novel feed-forward network that reconstructs a high-fidelity, animatable 3D Gaussian avatar from as few as a single image. To assist training, we create a large-scale synthetic motion dataset, VarenPoser, which features high-quality surface motions and diverse camera trajectories, as well as a synthetic appearance dataset, VarenTex, comprising realistic multi-view images generated through multi-view diffusion. While training only on synthetic datasets, 4DEquine achieves state-of-the-art performance on real-world APT36K and AiM datasets, demonstrating the superiority of 4DEquine and our new datasets for both geometry and appearance reconstruction. Comprehensive ablation studies validate the effectiveness of both the motion and appearance reconstruction network. Project page: https://luoxue-star.github.io/4DEquine_Project_Page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。