同步重建3D虚拟人与动作捕捉,提升精度与视觉质量
Better Together: Unified Motion Capture and 3D Avatar Reconstruction
- 用时序MLP联合优化骨骼运动与3D高斯体模型
- 人体关节误差降35%,手部关节误差降45%
- 适合虚拟人实时渲染与高保真动画应用
我们提出 Better Together,一种从多视角视频中同步完成人体姿态估计与真实感3D虚拟人重建的方法。以往方法分别处理这两项任务,而我们通过联合优化骨骼运动与可渲染3D身体模型,实现协同增益:不仅提高动作捕捉精度,还显著改善虚拟人实时渲染的视觉质量。为此,我们设计了一种基于个性化网格的可驱动3D高斯体模型,并采用时间依赖的MLP来优化运动序列,获得准确且时序一致的姿态估计。在极具挑战性的瑜伽动作数据集上评估,本方法在多视角人体姿态估计上达到当前最优水平,相较关键点基方法,身体关节约降低35%误差,手部关节约降低45%。同时,在多样复杂主体上,我们的方法将新视角合成的图像质量提升2dB PSNR。
原文摘要 · Abstract (English)
We present Better Together, a method that simultaneously solves the human pose estimation problem while reconstructing a photorealistic 3D human avatar from multi-view videos. While prior art usually solves these problems separately, we argue that joint optimization of skeletal motion with a 3D renderable body model brings synergistic effects, i.e. yields more precise motion capture and improved visual quality of real-time rendering of avatars. To achieve this, we introduce a novel animatable avatar with 3D Gaussians rigged on a personalized mesh and propose to optimize the motion sequence with time-dependent MLPs that provide accurate and temporally consistent pose estimates. We first evaluate our method on highly challenging yoga poses and demonstrate state-of-the-art accuracy on multi-view human pose estimation, reducing error by 35% on body joints and 45% on hand joints compared to keypoint-based methods. At the same time, our method significantly boosts the visual quality of animatable avatars (+2dB PSNR on novel view synthesis) on diverse challenging subjects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。