用预训练扩散模型实现高质量纹理生成,解决多视角一致性问题。
GenesisTex2: Stable, Consistent and High-Quality Text-to-Texture Generation

- 引入局部注意力重加权机制,聚焦跨视角空间相关区域。
- 在多个数据集上优于现有方法,保持细节清晰且无明显失真。
- 无需微调,兼容主流生成模型,适合快速部署使用。
大规模文本引导的图像扩散模型在文本到图像(T2I)生成任务中表现惊人,但将其应用于3D几何体纹理合成仍面临2D图像与3D表面纹理之间的领域差异挑战。早期基于投影-修复的方法虽保留了生成多样性,却常出现明显伪影和风格不一致。近期方法虽尝试改善一致性,但往往引入模糊、过饱和或过度平滑等问题。为此,我们提出一种新型文本到纹理合成框架,利用预训练扩散模型。首先,在自注意力层中引入局部注意力重加权机制,引导模型聚焦不同视角间的空间相关块,从而增强局部细节并保持跨视角一致性。此外,提出一种新颖的潜在空间融合管道,进一步确保多视角间的一致性,同时不显著牺牲多样性。实验表明,该方法在纹理一致性和视觉质量上显著优于现有最先进技术,且生成速度远超基于蒸馏的方法。重要的是,本框架无需额外训练或微调,可高度适配公开平台上的多种模型。
原文摘要 · Abstract (English)
Large-scale text-guided image diffusion models have shown astonishing results in text-to-image (T2I) generation. However, applying these models to synthesize textures for 3D geometries remains challenging due to the domain gap between 2D images and textures on a 3D surface. Early works that used a projecting-and-inpainting approach managed to preserve generation diversity but often resulted in noticeable artifacts and style inconsistencies. While recent methods have attempted to address these inconsistencies, they often introduce other issues, such as blurring, over-saturation, or over-smoothing. To overcome these challenges, we propose a novel text-to-texture synthesis framework that leverages pretrained diffusion models. We first introduce a local attention reweighing mechanism in the self-attention layers to guide the model in concentrating on spatial-correlated patches across different views, thereby enhancing local details while preserving cross-view consistency. Additionally, we propose a novel latent space merge pipeline, which further ensures consistency across different viewpoints without sacrificing too much diversity. Our method significantly outperforms existing state-of-the-art techniques regarding texture consistency and visual quality, while delivering results much faster than distillation-based methods. Importantly, our framework does not require additional training or fine-tuning, making it highly adaptable to a wide range of models available on public platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。