用自然语言控制亮度,实现个性化低光图像增强。
TSCnet: A Text-driven Semantic-level Controllable Framework for Customized Low-Light Image Enhancement
- 通过大模型理解文本指令,定位需增强的物体区域。
- 结合反射图生成精确掩码,实现语义级亮度可控调节。
- 支持复杂场景下自然语言交互,适合个性化图像修复应用。
基于深度学习的图像增强方法在降低噪声和提升低光照可见性方面表现优异,但通常采用一对一映射,缺乏个性化调整能力。为此,本文提出一种新的低光图像增强任务与框架——TSCnet,支持通过自然语言提示实现语义级、量化亮度控制。首先,利用大语言模型(LLM)解析自然语言指令,识别目标对象;随后,基于Retinex的推理分割(RRS)模块生成反射图像,输出精准的目标定位掩码;接着,文本驱动的亮度可控(TBC)模块依据光照图调整亮度;最后,自适应上下文补偿(ACC)模块融合多模态输入,调控条件扩散模型完成精细、无伪影的光影增强。在多个基准数据集上的实验表明,该框架显著提升了可见度,保持了自然色彩平衡,并增强了细节表现力。其强大的泛化能力使复杂开放世界环境中通过自然语言进行语义级照明调整成为可能。
原文摘要 · Abstract (English)
Deep learning-based image enhancement methods show significant advantages in reducing noise and improving visibility in low-light conditions. These methods are typically based on one-to-one mapping, where the model learns a direct transformation from low light to specific enhanced images. Therefore, these methods are inflexible as they do not allow highly personalized mapping, even though an individual's lighting preferences are inherently personalized. To overcome these limitations, we propose a new light enhancement task and a new framework that provides customized lighting control through prompt-driven, semantic-level, and quantitative brightness adjustments. The framework begins by leveraging a Large Language Model (LLM) to understand natural language prompts, enabling it to identify target objects for brightness adjustments. To localize these target objects, the Retinex-based Reasoning Segment (RRS) module generates precise target localization masks using reflection images. Subsequently, the Text-based Brightness Controllable (TBC) module adjusts brightness levels based on the generated illumination map. Finally, an Adaptive Contextual Compensation (ACC) module integrates multi-modal inputs and controls a conditional diffusion model to adjust the lighting, ensuring seamless and precise enhancements accurately. Experimental results on benchmark datasets demonstrate our framework's superior performance at increasing visibility, maintaining natural color balance, and amplifying fine details without creating artifacts. Furthermore, its robust generalization capabilities enable complex semantic-level lighting adjustments in diverse open-world environments through natural language interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。