用无监督微调打通低光增强与理解的桥梁,提升通用性和下游性能。
From Enhancement to Understanding: Build a Generalized Bridge for Low-light Vision via Semantically Consistent Unsupervised Fine-tuning
- 引入光照感知图像提示与循环注意力适配器,提升生成语义一致性。
- 在分类、检测、分割任务上均超越现有方法,实现零样本泛化。
- 适合需要低光场景通用处理能力的研究者与工程师。
低光增强与低光视觉理解传统上被分开处理。低光增强依赖物理或几何先验,泛化性受限,评估多关注视觉质量而非下游性能;低光理解因标注数据稀缺,主要采用任务特异的域适应,可扩展性差。为此,我们构建了低光增强与理解间的通用桥梁——通用增强理解框架(GEFU),提升泛化与可扩展性。针对低光退化多样性的挑战,利用预训练生成扩散模型优化图像,实现零样本泛化。在此基础上提出语义一致的无监督微调(SCUF):为克服文本提示局限,引入光照感知图像提示以显式引导生成,并设计循环注意力适配器最大化其语义潜力;为缓解无监督训练中的语义退化,提出标题与反射率一致性机制,学习高层语义与图像级空间语义。大量实验表明,所提方法在传统图像质量及分类、检测、语义分割等GEFU任务中均优于当前最先进方法。
原文摘要 · Abstract (English)
Low-level enhancement and high-level visual understanding in low-light vision have traditionally been treated separately. Low-light enhancement improves image quality for downstream tasks, but existing methods rely on physical or geometric priors, limiting generalization. Evaluation mainly focuses on visual quality rather than downstream performance. Low-light visual understanding, constrained by scarce labeled data, primarily uses task-specific domain adaptation, which lacks scalability. To address these challenges, we build a generalized bridge between low-light enhancement and low-light understanding, which we term Generalized Enhancement For Understanding (GEFU). This paradigm improves both generalization and scalability. To address the diverse causes of low-light degradation, we leverage pretrained generative diffusion models to optimize images, achieving zero-shot generalization performance. Building on this, we propose Semantically Consistent Unsupervised Fine-tuning (SCUF). Specifically, to overcome text prompt limitations, we introduce an illumination-aware image prompt to explicitly guide image generation and propose a cycle-attention adapter to maximize its semantic potential. To mitigate semantic degradation in unsupervised training, we propose caption and reflectance consistency to learn high-level semantics and image-level spatial semantics. Extensive experiments demonstrate that our proposed method outperforms current state-of-the-art methods in traditional image quality and GEFU tasks including classification, detection, and semantic segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。