arXiv:2512.02781cs.CVcs.GR2025-12

LumiX 一键生成场景的多张物理属性图,保持结构一致且真实。

LumiX: Structured and Coherent Text-to-Intrinsic Generation

  • 用共享查询机制保证各属性图结构一致
  • 联合生成漫反射、法线、深度等5张图,比当前最佳高23%对齐度
  • 适合需要高质量图像物理分解的研究与应用

我们提出 LumiX,一种用于连贯文本到内在属性生成的结构化扩散框架。在文本提示条件下,LumiX 联合生成一系列完整的内在图(如漫反射、辐照度、法线、深度和最终颜色),提供对底层场景的结构化且物理一致的描述。这得益于两项关键贡献:1)查询广播注意力,通过在每个自注意力块中跨所有图共享查询,确保结构一致性;2)张量 LoRA,一种基于张量的适配方法,参数高效地建模跨图关系,实现高效的联合训练。这些设计共同实现了稳定的联合扩散训练与多属性统一生成。实验表明,LumiX 生成结果具有一致性和物理合理性,在对齐度上比现有最优方法高23%,偏好评分达0.19(对比-0.41),且可在同一框架内完成图像条件下的内在分解。

原文摘要 · Abstract (English)

We present LumiX, a structured diffusion framework for coherent text-to-intrinsic generation. Conditioned on text prompts, LumiX jointly generates a comprehensive set of intrinsic maps (e.g., albedo, irradiance, normal, depth, and final color), providing a structured and physically consistent description of an underlying scene. This is enabled by two key contributions: 1) Query-Broadcast Attention, a mechanism that ensures structural consistency by sharing queries across all maps in each self-attention block. 2) Tensor LoRA, a tensor-based adaptation that parameter-efficiently models cross-map relations for efficient joint training. Together, these designs enable stable joint diffusion training and unified generation of multiple intrinsic properties. Experiments show that LumiX produces coherent and physically meaningful results, achieving 23% higher alignment and a better preference score (0.19 vs. -0.41) compared to the state of the art, and it can also perform image-conditioned intrinsic decomposition within the same framework.

图像生成扩散模型物理属性多图生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。