arXiv:2604.24877cs.CVcs.AI2026-04中稿 · ICLR

用自然语言控制图像光照,全开源可复现。

Learning Illumination Control in Diffusion Models

论文配图:Learning Illumination Control in Diffusion Models
图 1 · 摘自论文原文
  • 构建光照转换数据引擎,生成三元组训练数据。
  • 在SD 1.5、SDXL等模型上显著提升光照还原效果。
  • 适合需要可控光照生成的视觉创作与研究者使用。

控制图像光照对摄影和视觉内容创作至关重要。尽管闭源模型已展现出出色的光照控制能力,但开源替代方案要么需要深度图等大量控制输入,要么未公开数据与代码。本文提出一个完全开源且可复现的扩散模型光照控制流程。方法构建数据引擎,将光照良好的图像转换为包含低光照输入、自然语言光照指令和高光照输出的监督三元组。在该数据上微调扩散模型,显著优于基线SD 1.5、SDXL及FLUX.1-dev模型,在感知相似性、结构相似性和身份保留方面表现更优。本工作完全基于开源工具与公开数据构建,所有代码、数据与模型权重均已公开。

原文摘要 · Abstract (English)

Controlling illumination in images is essential for photography and visual content creation. While closed-source models have demonstrated impressive illumination control, open-source alternatives either require heavy control inputs like depth maps or do not release their data and code. We present a fully open-source and reproducible pipeline for learning illumination control in diffusion models. Our approach builds a data engine that transforms well-lit images into supervised training triplets consisting of a poorly-illuminated input image, a natural language lighting instruction, and a well-illuminated output image. We finetune a diffusion model on this data and demonstrate significant improvements over baseline SD 1.5, SDXL, and FLUX.1-dev models in perceptual similarity, structural similarity, and identity preservation. Our work provides a reproducible solution built entirely with open-source tools and publicly available data. We release all our code, data, and model weights publicly.

光照控制扩散模型生成图像开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。