无需相机元数据,用提示词生成逼真真实噪声图像
Diffusion-Based sRGB Real Noise Generation via Prompt-Driven Noise Representation Learning
- 通过提示词学习捕捉真实噪声特征,不依赖相机元数据
- 生成的噪声图像在多个基准数据集上有效用于去噪任务
- 适合需要真实噪声合成的图像处理与去噪研究者
在sRGB图像空间中去噪面临噪声高度变异的挑战。尽管端到端方法表现良好,但其在真实场景中的效果受限于真实噪声-清晰图像对的稀缺性,这类数据采集成本高且困难。为此,已有生成方法尝试从有限数据中合成逼真噪声图像,但通常需依赖训练和测试时的相机元数据,而元数据缺失或设备间不一致会限制其适用性。为此,我们提出一种名为提示驱动噪声生成(Prompt-Driven Noise Generation, PNG)的新框架。该模型可获取高维提示特征,捕获真实输入噪声的特性,并生成多种与输入噪声分布一致的逼真噪声图像。通过消除对显式相机元数据的依赖,显著提升了噪声合成的泛化能力和实用性。大量实验表明,本模型能有效生成逼真噪声图像,并成功应用于多个基准数据集的真实噪声去除任务。
原文摘要 · Abstract (English)
Denoising in the sRGB image space is challenging due to large noise variability. Although end-to-end methods perform well, their effectiveness in real-world scenarios is limited by the scarcity of real noisy-clean image pairs, which are expensive and difficult to collect. To address this limitation, several generative methods have been developed to synthesize realistic noisy images from limited data. These approaches often rely on camera metadata during both training and testing to synthesize real-world noise. However, the lack of metadata or inconsistencies between devices restricts their usability. Therefore, we propose a novel framework called Prompt-Driven Noise Generation (PNG). This model is capable of acquiring high-dimensional prompt features that capture the characteristics of real-world input noise and creating a variety of realistic noisy images consistent with the distribution of the input noise. By eliminating the dependency on explicit camera metadata, our approach significantly enhances the generalizability and applicability of noise synthesis. Comprehensive experiments reveal that our model effectively produces realistic noisy images and show the successful application of these generated images in removing real-world noise across various benchmark datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。