用RGB生成高保真RAW图像,少样本也能高效建模
RAW-Diffusion: RGB-Guided Diffusion Models for High-Fidelity RAW Image Generation
- 用RGB引导扩散模型,跨分辨率融合特征生成RAW图
- 仅25张训练样本即达顶尖性能,数据效率极高
- 适合缺乏大量RAW数据的视觉任务研究者使用
当前计算机视觉多聚焦于RGB数据,忽略更丰富的原始图像信息。而RAW图像在低光等挑战场景中更具表现力。然而,构建特定相机的RAW数据集成本高昂。为此,我们提出一种基于扩散模型的RGB引导方法,通过提取RGB输入特征,并将其融入反向扩散过程中的多尺度残差块,生成高保真RAW图像。该方法可高效构建相机专用的RAW数据集。在四个DSLR数据集上的实验表明,该方法达到当前最优水平。此外,其具备极强的数据效率,仅需25个训练样本即可实现优异效果。我们进一步将该方法应用于构建BDD100K-RAW与Cityscapes-RAW数据集,在原始图像上进行目标检测时显著减少了对真实RAW图像的需求。
原文摘要 · Abstract (English)
Current deep learning approaches in computer vision primarily focus on RGB data sacrificing information. In contrast, RAW images offer richer representation, which is crucial for precise recognition, particularly in challenging conditions like low-light environments. The resultant demand for comprehensive RAW image datasets contrasts with the labor-intensive process of creating specific datasets for individual sensors. To address this, we propose a novel diffusion-based method for generating RAW images guided by RGB images. Our approach integrates an RGB-guidance module for feature extraction from RGB inputs, then incorporates these features into the reverse diffusion process with RGB-guided residual blocks across various resolutions. This approach yields high-fidelity RAW images, enabling the creation of camera-specific RAW datasets. Our RGB2RAW experiments on four DSLR datasets demonstrate state-of-the-art performance. Moreover, RAW-Diffusion demonstrates exceptional data efficiency, achieving remarkable performance with as few as 25 training samples or even fewer. We extend our method to create BDD100K-RAW and Cityscapes-RAW datasets, revealing its effectiveness for object detection in RAW imagery, significantly reducing the amount of required RAW images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。