用局部线性混合形状实现手机端高保真虚拟人建模
High-Fidelity Mobile Avatars with Pruned Local Blendshapes

- 用小部位局部线性混合形状捕捉高斯属性的非线性变化
- 剔除变化小的高斯点对应混合形状,压缩模型至最小
- 无需预训练模型,支持移动端120帧2K渲染
本文提出一种从多视角视频重建高保真人体虚拟人的方法,可在移动设备上运行。现有基于高斯的全身建模方法虽质量高,但需大量计算生成姿态相关外观,难以部署于移动端。近期方法通过线性组合全局姿态特征与混合形状来建模姿态依赖的非线性高斯属性,虽可移动端运行,但细节损失明显。我们观察到身体局部区域内高斯点间高度相关,可用更少误差的线性方式建模。因此,采用小身体部位的局部线性混合形状以捕捉全局非线性变化。为进一步降低计算量和模型尺寸,提出移除属性变化微小的高斯点对应的混合形状,获得最小混合形状表示。方法为端到端训练,无需预训练模型。为适配多设备,使用WebGPU实现。实验表明,该方法可在移动端实现2K分辨率下120帧/秒的高质量人体渲染。
原文摘要 · Abstract (English)
We propose a method to reconstruct high-fidelity human avatars from multi-view video that can run on mobile devices. Many works can model high-quality Gaussian-based full-body avatars from multi-view video. However, these methods require heavy computation to obtain pose-dependent appearance, making deployment on mobile devices very difficult. Recent methods distill from pretrained models and model pose-dependent nonlinear Gaussian attributes by linearly combining global pose features with blendshapes. Although they can run on mobile devices, they suffer some loss of detail. We observe that nearby Gaussians are often highly correlated within a local region of the body, and can be linearly modeled with less error. Therefore, we use local linear blendshapes in small body parts to capture global nonlinear changes of Gaussian attributes. To further reduce computation and model size, we propose to remove blendshapes for Gaussians whose attributes change little, yielding a minimal blendshape representation. Our method is an end-to-end training method without a pretrained model. To make it run on multiple devices, we implement our method using WebGPU. Experiments show that our method can render high-quality human avatars with better details, and can reach 120 FPS at 2K resolution on mobile devices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。