arXiv:2409.12431cs.CVcs.AI2024-09中稿 · AAAI被引 11

用视觉引导提升纹理生成质量,解决模糊不一致问题

FlexiTex: Enhancing Texture Generation via Visual Guidance

  • 通过视觉引导模块增强文本提示,减少歧义并保留高频细节
  • 引入方向感知适配模块,自动生成相机姿态对应的提示
  • 显著改善纹理清晰度与一致性,适合真实场景应用

近期纹理生成方法得益于大规模文本到图像扩散模型的强大生成先验,取得了令人瞩目的成果。然而,抽象的文本提示在提供全局纹理或形状信息方面能力有限,导致生成结果出现模糊或不一致的图案。为此,我们提出FlexiTex,通过视觉引导嵌入丰富信息以生成高质量纹理。其核心是视觉引导增强模块,从视觉引导中融入更具体的信息,降低文本提示的歧义性,并保留高频细节。为进一步强化视觉引导,我们引入方向感知适配模块,根据不同的相机姿态自动设计方向提示,避免了双面性(Janus)问题,维持语义上的全局一致性。得益于视觉引导,FlexiTex在定量和定性上均表现出色,展现出推动真实世界纹理生成应用的潜力。

原文摘要 · Abstract (English)

Recent texture generation methods achieve impressive results due to the powerful generative prior they leverage from large-scale text-to-image diffusion models. However, abstract textual prompts are limited in providing global textural or shape information, which results in the texture generation methods producing blurry or inconsistent patterns. To tackle this, we present FlexiTex, embedding rich information via visual guidance to generate a high-quality texture. The core of FlexiTex is the Visual Guidance Enhancement module, which incorporates more specific information from visual guidance to reduce ambiguity in the text prompt and preserve high-frequency details. To further enhance the visual guidance, we introduce a Direction-Aware Adaptation module that automatically designs direction prompts based on different camera poses, avoiding the Janus problem and maintaining semantically global consistency. Benefiting from the visual guidance, FlexiTex produces quantitatively and qualitatively sound results, demonstrating its potential to advance texture generation for real-world applications.

纹理生成视觉引导扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。