无需相机参数,生成真实低光图像噪声,提升低光视觉任务性能
Towards a General-Purpose Zero-Shot Synthetic Low-Light Image and Video Pipeline
- 自监督训练的降解估计网络,估算物理噪声参数生成真实sRGB噪声
- 在噪声复现、视频增强、目标检测上分别提升24%、21%、62%性能
- 零样本适配,适用于多种设备和场景,适合低光数据增强研究者
低光照条件对人工与机器标注均构成挑战,导致低光图像与(尤其是)视频的机器理解研究匮乏。现有方法常将高质量数据集的标注迁移到合成的低光版本,但受限于不真实的噪声模型。本文提出新型降解估计网络(DEN),无需相机元数据即可合成真实的标准RGB(sRGB)噪声,通过自监督方式估计物理启发的噪声分布参数。该零样本方法可生成多样且真实的噪声特征,超越仅复现训练数据噪声的方法。我们在多种基于合成数据训练的任务上评估该管道:噪声复现、视频增强与目标检测,分别实现最高24%的KLD降低、21%的LPIPS下降与62%的AP$_{50-95}$提升。
原文摘要 · Abstract (English)
Low-light conditions pose significant challenges for both human and machine annotation. This in turn has led to a lack of research into machine understanding for low-light images and (in particular) videos. A common approach is to apply annotations obtained from high quality datasets to synthetically created low light versions. In addition, these approaches are often limited through the use of unrealistic noise models. In this paper, we propose a new Degradation Estimation Network (DEN), which synthetically generates realistic standard RGB (sRGB) noise without the requirement for camera metadata. This is achieved by estimating the parameters of physics-informed noise distributions, trained in a self-supervised manner. This zero-shot approach allows our method to generate synthetic noisy content with a diverse range of realistic noise characteristics, unlike other methods which focus on recreating the noise characteristics of the training data. We evaluate our proposed synthetic pipeline using various methods trained on its synthetic data for typical low-light tasks including synthetic noise replication, video enhancement, and object detection, showing improvements of up to 24\% KLD, 21\% LPIPS, and 62\% AP$_{50-95}$, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。