arXiv:2512.04267cs.CV2025-12被引 3

统一光照表示,让文本、图像等多模态光照信息可互通使用。

UniLight: A Unified Representation for Lighting

  • 构建共享嵌入空间,融合文本、图像、辐射度、环境图等光照模态。
  • 在光照检索、环境图生成和扩散图像合成中实现跨模态迁移性能提升。
  • 支持跨模态光照控制,适合需要灵活光照编辑的生成任务研究者。

光照对视觉外观影响显著,但图像中光照的理解与表示仍具挑战。现有表示方法如环境图、辐射度、球谐函数或文本互不兼容,限制了跨模态迁移。为此,我们提出UniLight,一种联合潜在空间的光照表示,将多种模态统一于共享嵌入中。通过对比学习训练文本、图像、辐射度和环境图的模态特定编码器,并引入辅助球谐函数预测任务以强化方向感知。大规模多模态数据管道支持三个任务的训练与评估:基于光照的检索、环境图生成,以及扩散模型中的光照控制。实验表明,该表示能捕捉一致且可迁移的光照特征,实现跨模态灵活操控。

原文摘要 · Abstract (English)

Lighting has a strong influence on visual appearance, yet understanding and representing lighting in images remains notoriously difficult. Various lighting representations exist, such as environment maps, irradiance, spherical harmonics, or text, but they are incompatible, which limits cross-modal transfer. We thus propose UniLight, a joint latent space as lighting representation, that unifies multiple modalities within a shared embedding. Modality-specific encoders for text, images, irradiance, and environment maps are trained contrastively to align their representations, with an auxiliary spherical-harmonics prediction task reinforcing directional understanding. Our multi-modal data pipeline enables large-scale training and evaluation across three tasks: lighting-based retrieval, environment-map generation, and lighting control in diffusion-based image synthesis. Experiments show that our representation captures consistent and transferable lighting features, enabling flexible manipulation across modalities.

光照表示多模态扩散模型生成任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。