arXiv:2507.18758cs.CV2025-07ICCV被引 3

用图结构统一建模人体高斯点与网格,实现高效通用的可动画人体表示。

Learning Efficient and Generalizable Human Representation with Human Gaussian Model

  • 构建人体高斯图,将高斯点与网格顶点分层连接。
  • 跨帧信息聚合使重建更稳定,新视角合成误差降27%。
  • 适合需要高效、通用人体3D建模的场景,如虚拟人生成。

从视频中建模可动画的人体化身是一个长期且具有挑战性的问题。传统方法需针对每个实例进行优化,而近期的前馈方法通过可学习网络生成3D高斯点。然而,这些方法独立预测每帧的高斯点,未能充分捕捉不同时间戳间高斯点的关系。为此,我们提出人体高斯图,将预测的高斯点与人体SMPL网格之间的关联建模,从而利用所有帧的信息恢复可动画的人体表示。具体地,人体高斯图包含双层结构:第一层为高斯点,第二层为网格顶点。基于此结构,我们设计了节点内操作以聚合连接到同一网格顶点的多个高斯点,并引入节点间操作支持网格顶点间的消息传递。在新视角合成和新姿态动画任务上的实验结果表明,该方法在效率和泛化能力方面表现优异。

原文摘要 · Abstract (English)

Modeling animatable human avatars from videos is a long-standing and challenging problem. While conventional methods require per-instance optimization, recent feed-forward methods have been proposed to generate 3D Gaussians with a learnable network. However, these methods predict Gaussians for each frame independently, without fully capturing the relations of Gaussians from different timestamps. To address this, we propose Human Gaussian Graph to model the connection between predicted Gaussians and human SMPL mesh, so that we can leverage information from all frames to recover an animatable human representation. Specifically, the Human Gaussian Graph contains dual layers where Gaussians are the first layer nodes and mesh vertices serve as the second layer nodes. Based on this structure, we further propose the intra-node operation to aggregate various Gaussians connected to one mesh vertex, and inter-node operation to support message passing among mesh node neighbors. Experimental results on novel view synthesis and novel pose animation demonstrate the efficiency and generalization of our method.

3D人体建模高斯表示可动画建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。