arXiv:2603.24086cs.CVcs.GR2026-03中稿 · IJCNN2026被引 1

无需训练,通过调整初始噪声实现文本图像生成的光照控制。

LGTM: Training-Free Light-Guided Text-to-Image Diffusion Model via Initial Noise Manipulation

  • 通过操控扩散过程的初始潜空间噪声实现光照引导。
  • 在不微调模型的前提下,光照一致性优于提示词基线方法。
  • 可无缝集成ControlNet,适合需要动态光照调节的场景。

扩散模型在条件文本到图像生成中表现出色,尤其在边缘、布局和深度等结构线索方面。然而,光照条件仍缺乏关注且难以控制。现有方法采用两阶段流程,在生成后重新打光,效率低;且依赖大规模数据微调与高计算量,限制了对新模型和任务的适应性。为此,我们提出一种无需训练的光照引导文生图扩散模型LGTM,通过操纵扩散过程的初始潜在噪声,结合文本提示与用户指定的光照方向来引导图像生成。通过对潜空间进行通道级分析,发现选择性地操控特定潜通道即可实现精细光照控制,无需微调或修改预训练模型。大量实验表明,该方法在光照一致性上优于基于提示的基线,同时保持图像质量和文本对齐。该方法为动态、用户可控的光照生成开辟了新可能,并能与ControlNet等模型无缝集成,适用于多种应用场景。

原文摘要 · Abstract (English)

Diffusion models have demonstrated high-quality performance in conditional text-to-image generation, particularly with structural cues such as edges, layouts, and depth. However, lighting conditions have received limited attention and remain difficult to control within the generative process. Existing methods handle lighting through a two-stage pipeline that relights images after generation, which is inefficient. Moreover, they rely on fine-tuning with large datasets and heavy computation, limiting their adaptability to new models and tasks. To address this, we propose a novel Training-Free Light-Guided Text-to-Image Diffusion Model via Initial Noise Manipulation (LGTM), which manipulates the initial latent noise of the diffusion process to guide image generation with text prompts and user-specified light directions. Through a channel-wise analysis of the latent space, we find that selectively manipulating latent channels enables fine-grained lighting control without fine-tuning or modifying the pre-trained model. Extensive experiments show that our method surpasses prompt-based baselines in lighting consistency, while preserving image quality and text alignment. This approach introduces new possibilities for dynamic, user-guided light control. Furthermore, it integrates seamlessly with models like ControlNet, demonstrating adaptability across diverse scenarios.

文生图扩散模型光照控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。