让模型主动学习光照变化,提升视觉识别鲁棒性。
Lighting-Aware Representation Learning under Controllable Lighting Variation

- 将光照变化作为显式训练信号,而非干扰因素。
- 在ImageNet等数据集上优于标准对比学习方法。
- 适用于复杂与简单光照场景,适合视觉任务通用增强。
光照变化仍是视觉表示学习的重大挑战,因其在环境间和环境中均引发显著外观变化。现有方法多通过数据增强使模型对光照变化保持不变,但未在学习过程中显式建模光照信息。受人类视觉理论启发,我们提出一种光照感知表示学习框架,将光照变化作为显式训练信号而非需抑制的干扰因素。该方法通过引入辅助目标,捕捉渲染场景中依赖光照的变化特征,使模型能联合学习保持语义一致、同时敏感于光照相关视觉结构的表示。我们在ImageNet、ExDark和PASCAL VOC上的图像分类与目标检测任务中评估了该模型。结果表明,相比标准对比学习基线,所提方法在相同架构和训练预算下持续提升下游性能。此外,该方法在监督学习框架及简化光照变化设置中也表现良好,显示出超越复杂光照场景的广泛适用性。这些结果表明其在复杂视觉环境及常规图像处理任务中增强模型鲁棒性与适应性的潜力。
原文摘要 · Abstract (English)
Variations in illumination remain a major challenge for visual representation learning, as they induce substantial appearance changes both across and within environments. While existing approaches typically address this issue through data augmentations that encourage models to become invariant to lighting changes, such strategies do not explicitly model lighting information during learning. Inspired by theories of human vision, we propose a lighting-aware representation learning framework that incorporates illumination variation as an explicit training signal rather than a nuisance factor to be suppressed. Our method extends contrastive learning by introducing an auxiliary objective that captures illumination-dependent variation in rendered scenes, enabling the model to jointly learn representations that preserve semantic consistency while remaining sensitive to lighting-dependent visual structure. We evaluate the proposed model on image classification and object detection tasks across the ImageNet, ExDark, and PASCAL VOC benchmarks. Results demonstrate that the proposed lighting-aware training consistently improves downstream performance over standard contrastive learning baselines, while maintaining the same architecture and training budget. Furthermore, our approach shows promising performance in supervised learning frameworks and under settings involving simpler lighting variation, suggesting broad applicability beyond complex illumination scenarios. These results indicate its potential to enhance model robustness and adaptability in complex visual environments as well as in more conventional image processing tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。