arXiv:2604.01447cs.CVcs.AI2026-04被引 1

用更优人体模型替代复杂网络,提升3D avatar重建质量

Better Rigs, Not Bigger Networks: A Body Model Ablation for Gaussian Avatars

论文配图:Better Rigs, Not Bigger Networks: A Body Model Ablation for Gaussian Avatars
图 1 · 摘自论文原文
  • 用SAM-3D-Body估算的MHR人体骨架替代SMPL,无需学习变形
  • 在PeopleSnapshot和ZJU-MoCap上达到最高PSNR,LPIPS与SSIM领先
  • 适合关注人体建模效率与质量的3D视觉研究者

近期基于SMPL的3D高斯点云方法虽实现优异视觉保真度,但训练架构持续复杂化。本文证明:大量复杂性实属冗余——仅用由SAM-3D-Body估计的Momentum Human Rig(MHR)替换SMPL,采用无学习形变或姿态依赖修正的极简流程,便在PeopleSnapshot与ZJU-MoCap数据集上达到最高报告的PSNR,且在LPIPS与SSIM上表现竞争或更优。为分离姿态估计质量与人体模型表达能力的影响,我们进行两项受控消融:将SAM-3D-Body网格转换为SMPL-X,以及将原始数据集中SMPL姿态转为MHR并在相同条件下重训练。结果表明,人体模型表达力是虚拟形象重建中的主要瓶颈,网格表示能力与姿态估计精度均对整体性能有显著贡献。

原文摘要 · Abstract (English)

Recent 3D Gaussian splatting methods built atop SMPL achieve remarkable visual fidelity while continually increasing the complexity of the overall training architecture. We demonstrate that much of this complexity is unnecessary: by replacing SMPL with the Momentum Human Rig (MHR), estimated via SAM-3D-Body, a minimal pipeline with no learned deformations or pose-dependent corrections achieves the highest reported PSNR and competitive or superior LPIPS and SSIM on PeopleSnapshot and ZJU-MoCap. To disentangle pose estimation quality from body model representational capacity, we perform two controlled ablations: translating SAM-3D-Body meshes to SMPL-X, and translating the original dataset's SMPL poses into MHR both retrained under identical conditions. These ablations confirm that body model expressiveness has been a primary bottleneck in avatar reconstruction, with both mesh representational capacity and pose estimation quality contributing meaningfully to the full pipeline's gains.

3D重建人体建模高斯点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。