arXiv:2504.12909cs.GRcs.CV2025-04CVPR被引 17

用空间分布MLP实现高保真人体动画,实时渲染且细节丰富。

Real-time High-fidelity Gaussian Human Avatars with Position-based Interpolation of Spatially Distributed MLPs

  • MLP按身体位置分布,通过距离加权插值生成高斯属性。
  • 引入偏移基函数让细节变化更自由,保留高频信号特征。
  • 控制点约束表面分布,新姿态泛化更好,适合实时应用。

现有方法在多视角视频中重建高斯人体模型时,或难以捕捉姿态依赖的外观细节,或依赖计算量大的神经网络导致渲染速度无法实现实时。本文提出一种新型高斯人体模型,可在保持高保真姿态依赖外观的同时实现实时渲染。模型采用分布在人体不同位置的MLP,每个高斯的参数通过其邻近MLP输出按距离插值得到。为避免插值导致属性平滑失真,为每个高斯定义一组高斯偏移基,由线性组合表示相对于中性状态的属性偏移。令MLP输出对应基的系数,使系数平滑变化而基自由学习,从而能捕捉高频空间信号。进一步使用控制点约束高斯仅分布在表层而非体内,提升新姿态下的泛化能力。相比当前最优方法,本方法在新视角和新姿态下显著提升渲染速度并获得更精细的外观质量。

原文摘要 · Abstract (English)

Many works have succeeded in reconstructing Gaussian human avatars from multi-view videos. However, they either struggle to capture pose-dependent appearance details with a single MLP, or rely on a computationally intensive neural network to reconstruct high-fidelity appearance but with rendering performance degraded to non-real-time. We propose a novel Gaussian human avatar representation that can reconstruct high-fidelity pose-dependence appearance with details and meanwhile can be rendered in real time. Our Gaussian avatar is empowered by spatially distributed MLPs which are explicitly located on different positions on human body. The parameters stored in each Gaussian are obtained by interpolating from the outputs of its nearby MLPs based on their distances. To avoid undesired smooth Gaussian property changing during interpolation, for each Gaussian we define a set of Gaussian offset basis, and a linear combination of basis represents the Gaussian property offsets relative to the neutral properties. Then we propose to let the MLPs output a set of coefficients corresponding to the basis. In this way, although Gaussian coefficients are derived from interpolation and change smoothly, the Gaussian offset basis is learned freely without constraints. The smoothly varying coefficients combined with freely learned basis can still produce distinctly different Gaussian property offsets, allowing the ability to learn high-frequency spatial signals. We further use control points to constrain the Gaussians distributed on a surface layer rather than allowing them to be irregularly distributed inside the body, to help the human avatar generalize better when animated under novel poses. Compared to the state-of-the-art method, our method achieves better appearance quality with finer details while the rendering speed is significantly faster under novel views and novel poses.

人体建模实时渲染高斯表示姿态细节

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。