arXiv:2506.14549cs.CV2025-06NeurIPS被引 8

DreamLight实现自然光影融合,支持图文双模式图像重照明。

DreamLight: Towards Harmonious and Consistent Image Relighting

  • 统一输入格式,利用预训练扩散模型语义先验生成真实光照效果。
  • 提出方向引导光适配器,精准控制前景与背景的光照一致性。
  • 新增频谱前景修复模块,提升主体与背景的视觉一致性,适合内容创作使用。

本文提出DreamLight模型,实现通用图像重照明,可无缝将主体合成至新背景中,保持光照与色彩基调的美学统一。背景可通过自然图像(图像驱动重照明)或任意文本提示生成(文本驱动重照明)。现有研究多聚焦图像驱动场景,对文本驱动探索不足。部分方法依赖环境图设计复杂解耦流程,需高昂数据成本;另一些则将任务视为图像翻译,采用自编码器进行像素级变换,虽取得较好融合效果,但难以生成真实的前景-背景光照交互。为此,我们重构输入数据为统一格式,借助预训练扩散模型的语义先验生成自然结果。提出位置引导光适配器(PGLA),将背景多方向光照信息压缩为光查询嵌入,并通过方向偏置掩码注意力调控前景。此外,设计后处理模块频谱前景修复器(SFF),自适应重组主体与重照明背景的不同频率成分,增强外观一致性。大量对比实验与用户研究证明,DreamLight在重照明性能上表现卓越。

原文摘要 · Abstract (English)

We introduce a model named DreamLight for universal image relighting in this work, which can seamlessly composite subjects into a new background while maintaining aesthetic uniformity in terms of lighting and color tone. The background can be specified by natural images (image-based relighting) or generated from unlimited text prompts (text-based relighting). Existing studies primarily focus on image-based relighting, while with scant exploration into text-based scenarios. Some works employ intricate disentanglement pipeline designs relying on environment maps to provide relevant information, which grapples with the expensive data cost required for intrinsic decomposition and light source. Other methods take this task as an image translation problem and perform pixel-level transformation with autoencoder architecture. While these methods have achieved decent harmonization effects, they struggle to generate realistic and natural light interaction effects between the foreground and background. To alleviate these challenges, we reorganize the input data into a unified format and leverage the semantic prior provided by the pretrained diffusion model to facilitate the generation of natural results. Moreover, we propose a Position-Guided Light Adapter (PGLA) that condenses light information from different directions in the background into designed light query embeddings, and modulates the foreground with direction-biased masked attention. In addition, we present a post-processing module named Spectral Foreground Fixer (SFF) to adaptively reorganize different frequency components of subject and relighted background, which helps enhance the consistency of foreground appearance. Extensive comparisons and user study demonstrate that our DreamLight achieves remarkable relighting performance.

图像重照明扩散模型光照一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。