arXiv:2503.19011cs.CV2025-03ICCV被引 21

用3D感知的注意力机制,让纹理生成更真实无瑕疵。

RomanTex: Decoupling 3D-aware Rotary Positional Embedded Multi-Attention Network for Texture Synthesis

  • 将3D几何与2D扩散模型结合,用新型旋转位置编码对齐多视角信息。
  • 在多个数据集上实现最高纹理质量,减少拼接和鬼影伪影。
  • 适合需要高质量3D纹理的建模、游戏与影视工业应用。

为现有几何体绘制纹理是3D资产生成中的关键但耗时任务。当前文本到图像(T2I)模型虽推动了纹理生成进展,但多数方法先在2D空间生成图像,再通过纹理烘焙映射到UV贴图,常因多视角图像不一致导致接缝与鬼影伪影。3D-based方法虽可缓解此问题,却忽略2D扩散模型先验,难以应用于真实物体。为此,我们提出RomanTex,一种基于多视角的纹理生成框架,通过新型3D-aware Rotary Positional Embedding,将多注意力网络与3D表示结合,并在注意力模块中引入解耦设计,增强对图像到纹理任务的鲁棒性,支持语义正确的背面合成。此外,引入与几何相关的无分类器引导(CFG)机制,进一步提升几何与图像的一致性。定量与定性评估及用户研究均表明,该方法在纹理质量与一致性方面达到当前最优水平。

原文摘要 · Abstract (English)

Painting textures for existing geometries is a critical yet labor-intensive process in 3D asset generation. Recent advancements in text-to-image (T2I) models have led to significant progress in texture generation. Most existing research approaches this task by first generating images in 2D spaces using image diffusion models, followed by a texture baking process to achieve UV texture. However, these methods often struggle to produce high-quality textures due to inconsistencies among the generated multi-view images, resulting in seams and ghosting artifacts. In contrast, 3D-based texture synthesis methods aim to address these inconsistencies, but they often neglect 2D diffusion model priors, making them challenging to apply to real-world objects To overcome these limitations, we propose RomanTex, a multiview-based texture generation framework that integrates a multi-attention network with an underlying 3D representation, facilitated by our novel 3D-aware Rotary Positional Embedding. Additionally, we incorporate a decoupling characteristic in the multi-attention block to enhance the model's robustness in image-to-texture task, enabling semantically-correct back-view synthesis. Furthermore, we introduce a geometry-related Classifier-Free Guidance (CFG) mechanism to further improve the alignment with both geometries and images. Quantitative and qualitative evaluations, along with comprehensive user studies, demonstrate that our method achieves state-of-the-art results in texture quality and consistency.

纹理生成3D-aware扩散模型多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。