arXiv:2504.01008cs.CVcs.AI2025-04NeurIPS被引 16

用文本生成可重打光的物理渲染图,提升内容创作灵活性

IntrinsiX: High-Quality PBR Generation using Image Priors

  • 分步训练各材质属性模型,通过跨内在注意力融合信息
  • 生成细节清晰的PBR图,对生成图像的分解性能显著优于现有方法
  • 适合需要重打光、编辑纹理的图形应用开发者

我们提出IntrinsiX,一种从文本描述生成高质量固有图像的新方法。与现有文本到图像模型输出中包含固定场景光照不同,本方法预测物理基础渲染(PBR)贴图,使生成结果可用于需重打光、编辑和纹理生成的核心图形应用场景。为训练生成器,我们利用强图像先验,分别预训练各PBR材质分量(反照率、粗糙度、金属度、法线)模型,并通过新型跨内在注意力机制,一致地拼接键值特征,实现各输出模态间信息交换,获得语义一致的PBR预测。为约束每种固有成分,我们设计渲染损失,在图像空间提供信号,促进输出BRDF属性中的精细细节。结果表明,该方法生成的固有图像具有丰富细节和强泛化能力,显著优于现有用于生成图像的固有图像分解方法。最后,我们展示了重打光、编辑及文本条件下的室级PBR纹理生成等应用。

原文摘要 · Abstract (English)

We introduce IntrinsiX, a novel method that generates high-quality intrinsic images from text description. In contrast to existing text-to-image models whose outputs contain baked-in scene lighting, our approach predicts physically-based rendering (PBR) maps. This enables the generated outputs to be used for content creation scenarios in core graphics applications that facilitate re-lighting, editing, and texture generation tasks. In order to train our generator, we exploit strong image priors, and pre-train separate models for each PBR material component (albedo, roughness, metallic, normals). We then align these models with a new cross-intrinsic attention formulation that concatenates key and value features in a consistent fashion. This allows us to exchange information between each output modality and to obtain semantically coherent PBR predictions. To ground each intrinsic component, we propose a rendering loss which provides image-space signals to constrain the model, thus facilitating sharp details also in the output BRDF properties. Our results demonstrate detailed intrinsic generation with strong generalization capabilities that outperforms existing intrinsic image decomposition methods used with generated images by a significant margin. Finally, we show a series of applications, including re-lighting, editing, and text-conditioned room-scale PBR texture generation.

图像生成PBR文本生成图形应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。