arXiv:2512.17489cs.CV2025-12中稿 · CVPR被引 1

让AI根据单张图学习光照风格,精准控制生成图像的光线效果。

LumiCtrl : Learning Illuminant Prompts for Lighting Control in Personalized Text-to-Image Models

  • 用物理光照增强和普朗克轨迹生成多种光照变体,辅助模型学习。
  • 通过边缘引导提示解耦,确保提示只关注光照而非物体结构。
  • 掩码重建损失实现前景光照精准控制,背景自适应融合,适合设计师使用。

文本到图像(T2I)模型在创意图像生成方面取得了显著进展,但仍缺乏对场景光源的精确控制,而光源是内容设计师调节生成图像视觉美感的关键因素。本文提出一种名为 LumiCtrl 的光照个性化方法,该方法基于单张物体图像学习光照提示。LumiCtrl 包含三个组件:(a)结合物理光照增强与普朗克轨迹,在标准光源下生成微调变体;(b)采用冻结的 ControlNet 进行边缘引导提示解耦,确保提示聚焦于光照而非结构;(c)使用掩码重建损失,聚焦于前景物体的光照学习,同时允许背景上下文自适应,实现所谓的上下文光适应。定性和定量对比表明,相较于现有基线方法,LumiCtrl 在光照保真度、美学质量及场景一致性方面均有显著提升。人类偏好研究进一步证实用户更倾向 LumiCtrl 生成结果。

原文摘要 · Abstract (English)

Text-to-image (T2I) models have demonstrated remarkable progress in creative image generation, yet they still lack precise control over scene illuminants which is a crucial factor for content designers to manipulate visual aesthetics of generated images. In this paper, we present an illuminant personalization method named LumiCtrl that learns illuminant prompt given single image of the object. LumiCtrl consists of three components: given an image of the object, our method apply (a) physics-based illuminant augmentation along with Planckian locus to create fine-tuning variants under standard illuminants; (b) Edge-Guided Prompt Disentanglement using frozen ControlNet to ensure prompts focus on illumination, not the structure; and (c) a Masked Reconstruction Loss that focuses learning on foreground object while allowing background to adapt contextually which enables what we call Contextual Light Adaptation. We qualitatively and quantitatively compare LumiCtrl against other T2I customization methods. The results show that LumiCtrl achieves significantly better illuminant fidelity, aesthetic quality, and scene coherence compared to existing baselines. A human preference study further confirms the strong user preference for LumiCtrl generations.

光照控制文本生成个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。