arXiv:2410.06985cs.CVcs.GR2024-10被引 5

让扩散模型一次生成多视角一致的PBR贴图,无需后期拼接。

Jointly Generating Multi-view Consistent PBR Textures using Collaborative Control

  • 用协同控制直接建模PBR贴图概率分布,支持法线、凹凸等多通道输出。
  • 在多视角一致性测试中显著优于现有方法,避免了复杂的融合步骤。
  • 适合游戏/影视资产生成,尤其需要高保真材质对齐的场景。

多视角一致性仍是图像扩散模型的挑战。即使在文本到贴图任务中已知精确几何对应关系,许多方法仍无法在不同视角间生成对齐结果,需借助复杂融合流程将输出映射回原始网格。本文聚焦于协同控制工作流在物理基础渲染(PBR)文本到贴图任务中的应用。协同控制直接建模包括法线、凹凸贴图在内的完整PBR图像概率分布,据我们所知,是首个能直接输出全栈PBR贴图的扩散模型。本文探讨了实现多视角一致性的设计决策,并通过消融实验和实际应用验证了该方法的有效性。

原文摘要 · Abstract (English)

Multi-view consistency remains a challenge for image diffusion models. Even within the Text-to-Texture problem, where perfect geometric correspondences are known a priori, many methods fail to yield aligned predictions across views, necessitating non-trivial fusion methods to incorporate the results onto the original mesh. We explore this issue for a Collaborative Control workflow specifically in PBR Text-to-Texture. Collaborative Control directly models PBR image probability distributions, including normal bump maps; to our knowledge, the only diffusion model to directly output full PBR stacks. We discuss the design decisions involved in making this model multi-view consistent, and demonstrate the effectiveness of our approach in ablation studies, as well as practical applications.

PBR贴图扩散模型多视角生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。