arXiv:2605.25778cs.CV2026-05

无需3D结构信息,直接从人脸图像生成可编辑的纹理图

OMGTex: One-stage Multi-style Facial Texture Reconstruction without Geometry Guidance

论文配图:OMGTex: One-stage Multi-style Facial Texture Reconstruction without Geometry Guidance
图 1 · 摘自论文原文
  • 不依赖3D几何先验,端到端生成面部UV纹理
  • 通过梯度引导修正结构错位,实现风格一致的纹理重建
  • 支持区域级编辑,适合需要风格化人脸的应用

我们提出OMGTex,一种基于扩散模型的端到端框架,可从多风格人脸图像中重建高质量且可编辑的面部UV纹理。现有方法存在两大局限:一是依赖难以准确估计的3D几何先验,尤其在遮挡或风格化场景下脆弱;二是缺乏语义解耦,阻碍区域级纹理编辑与风格迁移。本工作同时解决上述问题,提出无几何引导的流程,直接将2D人脸图像映射为可编辑的UV纹理。核心创新包括:一、引入推理时梯度引导的精修策略,显式纠正扩散生成中的结构不一致性;二、利用扩散模型固有的语义分布能力,设计新型训练范式,增强语义感知编辑能力。此外,为缓解多风格纹理重建数据稀缺问题,构建了首个多风格配对纹理重建数据集CANVAS,覆盖真实与多样化风格域。据我们所知,OMGTex是首个实现跨多样领域鲁棒、风格一致且可编辑面部纹理重建的无几何引导框架,在多个面部纹理基准上达到最先进性能。

原文摘要 · Abstract (English)

We propose OMGTex, an end-to-end diffusion-based framework for reconstructing high-quality and editable facial UV textures from multi-style facial images. Existing texture reconstruction methods face two major limitations: (1) Fragility due to reliance on 3D geometry priors, which are difficult to estimate accurately, especially under facial occlusions or in stylized domains; and (2) A lack of semantic disentanglement, inhibiting region-specific texture editing and style transfer. Our work addresses both challenges simultaneously. Our core innovation is a geometry-free pipeline that directly maps a 2D face image to its corresponding editable UV texture. We introduce two key techniques: First, to address the challenge of UV misalignment common in diffusion generation, we introduce a gradient-guided refinement strategy at inference time, which explicitly corrects structural consistency. Second, we leverage the inherent semantic distribution capability of diffusion models and design a novel training paradigm to enhance this tendency, enabling semantic-aware editing of facial texture. Furthermore, to address the data scarcity in multi-style texture reconstruction, we construct CANVAS, the first comprehensive paired texture reconstruction dataset covering realistic and diverse stylized domains. To the best of our knowledge, OMGTex is the first geometry-free inference framework that achieves robust, style-consistent, and editable facial texture reconstruction across diverse domains. Our method achieves state-of-the-art performance on multiple facial texture benchmarks.

面部纹理扩散模型无几何引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。