arXiv:2605.14594cs.CVcs.GR2026-05

让单图生成的3D头像统一拓扑,便于工业级动画使用。

TOPOS: High-Fidelity and Efficient Industry-Grade 3D Head Generation

论文配图:TOPOS: High-Fidelity and Efficient Industry-Grade 3D Head Generation
图 1 · 摘自论文原文
  • 用可学习的变分自编码器将不同拓扑头像转为固定工业拓扑。
  • 生成头像在顶点级别对齐,支持角色复用和动画绑定。
  • 适合影视游戏行业需批量生成一致3D人头的场景。

高保真3D头像生成在影视、动画和游戏产业中至关重要。工业管线中,工作室通常要求所有头像资产采用固定的参考拓扑,以满足生产级绑定、蒙皮与动画需求。本文提出TOPOS框架,面向单图条件下的3D头像生成,联合恢复几何与外观,并遵循行业标准拓扑。与产生不一致拓扑和大量顶点的一般3D生成模型不同,TOPOS生成具有固定、制片风格拓扑的头像网格,实现所有生成头像间的顶点级对应。为建模统一拓扑下的头像,我们提出新型变分自编码器TOPOS-VAE,借鉴多模态大语言模型思想,利用Perceiver Resampler将来自多种拓扑头像点云转换为目标参考拓扑。基于TOPOS-VAE的结构化隐空间,我们训练了修正流变换器TOPOS-DiT,高效地从单张图像生成高保真头像网格。我们还提出TOPOS-Texture,一个端到端模块,通过微调多模态图像生成模型,从同一肖像图像生成可光照重演的UV纹理图。生成的纹理与底层网格几何空间对齐,并忠实保留高频外观细节。大量实验表明,TOPOS在3D头像生成上达到领先性能,超越经典人脸重建方法与通用3D物体生成模型,凸显其在数字人创作中的有效性。

原文摘要 · Abstract (English)

High-fidelity 3D head generation plays a crucial role in the film, animation and video game industries. In industrial pipelines, studios typically enforce a fixed reference topology across all head assets, as such a clean and uniform topology is a prerequisite for production-level rigging, skinning and animation. In this paper, we present TOPOS, a framework tailored for single image conditioned 3D head generation that jointly recovers geometry and appearance under such an industry-standard topology. In contrast to general 3D generative models which produce triangle meshes with inconsistent topology and numerous vertices, hindering semantic correspondence and asset-level reuse, TOPOS generates head meshes with a fixed, studio-style topology, enabling consistent vertex-level correspondence across all generated heads. To model heads under this unified topology, we proposed a novel variational autoencoder structure, termed TOPOS-VAE. Inspired by multi-model large language models (MLLMs), our TOPOS-VAE leverages the Perceiver Resampler to convert input pointclouds sampled from head meshes of diverse topologies into the target reference topology. Building upon TOPOS-VAE's structured latent space, we train a rectified flow transformer, TOPOS-DiT, to efficiently generate high-fidelity head meshes from a single image. We further present TOPOS-Texture, an end-to-end module that produces relightable UV texture maps from the same portrait image via fine-tuning a multimodal image generative model. The generated textures are spatially aligned with the underlying mesh geometry and faithfully preserve high-frequency appearance details. Extensive experiments demonstrate that TOPOS achieves state-of-the-art performance on 3D head generation, surpassing both classical face reconstruction methods and general 3D object generative models, highlighting its effectiveness for digital human creation.

3D生成数字人拓扑统一工业级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。