arXiv:2604.18583cs.CV2026-04

让手机能跑高精度可动虚拟人,速度超180帧/秒。

MUA: Mobile Ultra-detailed Animatable Avatars

论文配图:MUA: Mobile Ultra-detailed Animatable Avatars
图 1 · 摘自论文原文
  • 用小波引导多层级分解+纹理低秩因子化,压缩模型体积。
  • 计算量降2000倍、模型缩小10倍,仍保持逼真动态细节。
  • 适合移动端沉浸式应用,手机端实现实时运行。

构建逼真且可动画化的全身数字人仍是计算机图形与视觉领域的长期挑战。现有方法主要沿两个方向发展:提升动态几何与外观保真度,或降低计算开销以适配资源受限平台(如VR头显)。但二者难以兼顾:超高质量虚拟人需依赖服务器级GPU,而轻量化模型常缺乏表面动态、细节不足且存在明显伪影。为此,本文提出一种新型可动画化虚拟人表示——小波引导的多层级空间因子化混合变形(Wavelet-guided Multi-level Spatial Factorized Blendshapes),并设计相应蒸馏流程,将预训练高保真教师模型中的运动感知服装动态与精细外观细节迁移到紧凑高效的表示中。通过在纹理空间结合多层级小波谱分解与低秩结构因子化,本方法实现较原始教师模型降低2000倍计算成本、模型尺寸缩小10倍,同时保持与教师模型高度接近的视觉合理性与细节表现。大量对比实验表明,本方法显著优于现有移动端虚拟人方案,在渲染质量上达到甚至超越多数仅能在服务器运行的方法。重要的是,该表示极大提升了高保真虚拟人在沉浸式应用中的实用性,在桌面电脑上实现超过180 FPS,且在独立式Meta Quest 3设备上实现24 FPS的原生实时性能。

原文摘要 · Abstract (English)

Building photorealistic, animatable full-body digital humans remains a longstanding challenge in computer graphics and vision. Recent advances in animatable avatar modeling have largely progressed along two directions: improving the fidelity of dynamic geometry and appearance, or reducing computational complexity to enable deployment on resource-constrained platforms, e.g., VR headsets. However, existing approaches fail to achieve both goals simultaneously: Ultra-high-fidelity avatars typically require substantial computation on server-class GPUs, whereas lightweight avatars often suffer from limited surface dynamics, reduced appearance details, and noticeable artifacts. To bridge this gap, we propose a novel animatable avatar representation, termed Wavelet-guided Multi-level Spatial Factorized Blendshapes, and a corresponding distillation pipeline that transfers motion-aware clothing dynamics and fine-grained appearance details from a pre-trained ultra-high-quality avatar model into a compact, efficient representation. By coupling multi-level wavelet spectral decomposition with low-rank structural factorization in texture space, our method achieves up to 2000X lower computational cost and a 10X smaller model size than the original high-quality teacher avatar model, while preserving visually plausible dynamics and appearance details closely resemble those of the teacher model. Extensive comparisons with state-of-the-art methods show that our approach significantly outperforms existing avatar approaches designed for mobile settings and achieves comparable or superior rendering quality to most approaches that can only run on servers. Importantly, our representation substantially improves the practicality of high-fidelity avatars for immersive applications, achieving over 180 FPS on a desktop PC and real-time native on-device performance at 24 FPS on a standalone Meta Quest 3.

虚拟人轻量化移动端动画生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。