arXiv:2603.01506cs.CV2026-03

单图0.2秒生成可动画3D头像,支持多精度适配不同设备

OMG-Avatar: One-shot Multi-LOD Gaussian Head Avatar

  • 用多层级高斯表示+变换器提取全局特征,投影采样获取局部细节
  • 0.2秒完成重建,相比现有方法在表情重演和质量上均更优
  • 适合实时虚拟形象、元宇宙交互等对速度与精度要求高的场景

我们提出OMG-Avatar,一种新颖的单图快速3D头像重建方法,可在0.2秒内基于单张图像生成可动画的3D头像。该方法采用多层级细节(Multi-LOD)高斯表示,统一模型支持不同硬件性能与推理速度需求。为同时捕捉全局与局部面部特征,采用基于变换器的架构提取全局特征,并通过投影采样获取局部特征;二者在深度缓冲引导下融合,确保遮挡合理性。进一步引入粗到精学习范式,实现层级细节感知。针对3DMM在肩部等非头部区域建模能力不足的问题,提出多区域分解策略,分别预测头与肩区域后通过跨区域组合融合。大量实验表明,OMG-Avatar在重建质量、表情重演表现及计算效率方面均优于当前最先进方法。

原文摘要 · Abstract (English)

We propose OMG-Avatar, a novel One-shot method that leverages a Multi-LOD (Level-of-Detail) Gaussian representation for animatable 3D head reconstruction from a single image in 0.2s. Our method enables LOD head avatar modeling using a unified model that accommodates diverse hardware capabilities and inference speed requirements. To capture both global and local facial characteristics, we employ a transformer-based architecture for global feature extraction and projection-based sampling for local feature acquisition. These features are effectively fused under the guidance of a depth buffer, ensuring occlusion plausibility. We further introduce a coarse-to-fine learning paradigm to support Level-of-Detail functionality and enhance the perception of hierarchical details. To address the limitations of 3DMMs in modeling non-head regions such as the shoulders, we introduce a multi-region decomposition scheme in which the head and shoulders are predicted separately and then integrated through cross-region combination. Extensive experiments demonstrate that OMG-Avatar outperforms state-of-the-art methods in reconstruction quality, reenactment performance, and computational efficiency. The project homepage is https://human3daigc.github.io/OMGAvatar_project_page/ .

3D头像单图重建高斯表示实时渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。