一个模型搞定多种相机的RAW图像转换,提升画质与效率
MERIT: Multi-domain Efficient RAW Image Translation
- 用统一模型实现任意相机间RAW图像互转
- 噪声统计对齐使生成图像质量提升5.56 dB
- 适合多相机图像处理与跨设备视觉任务
不同相机传感器捕获的RAW图像因光谱响应、噪声特性和色调行为差异,存在显著域偏移,难以直接用于下游计算机视觉任务。以往方法需为每对源-目标相机训练专用的RAW-to-RAW翻译器,难以扩展到包含多种商业相机的实际场景。本文提出MERIT,首个统一的多域RAW图像翻译框架,仅用一个模型即可实现任意相机域间的翻译。针对域间噪声差异,提出传感器感知噪声建模损失,显式对齐生成图像与目标域的信号依赖噪声统计;并通过条件多尺度大核注意力模块增强生成器的上下文与传感器感知特征建模能力。为促进标准化评估,引入MDRAW数据集,包含五种不同相机传感器在广泛场景下的配对与非配对RAW图像。大量实验表明,MERIT在质量上优于先前模型(提升5.56 dB),且可扩展性更强(训练迭代减少80%)。
原文摘要 · Abstract (English)
RAW images captured by different camera sensors exhibit substantial domain shifts due to varying spectral responses, noise characteristics, and tone behaviors, complicating their direct use in downstream computer vision tasks. Prior methods address this problem by training domain-specific RAW-to-RAW translators for each source-target pair, but such approaches do not scale to real-world scenarios involving multiple types of commercial cameras. In this work, we introduce MERIT, the first unified framework for multi-domain RAW image translation, which leverages a single model to perform translations across arbitrary camera domains. To address domain-specific noise discrepancies, we propose a sensor-aware noise modeling loss that explicitly aligns the signal-dependent noise statistics of the generated images with those of the target domain. We further enhance the generator with a conditional multi-scale large kernel attention module for improved context and sensor-aware feature modeling. To facilitate standardized evaluation, we introduce MDRAW, the first dataset tailored for multi-domain RAW image translation, comprising both paired and unpaired RAW captures from five diverse camera sensors across a wide range of scenes. Extensive experiments demonstrate that MERIT outperforms prior models in both quality (5.56 dB improvement) and scalability (80% reduction in training iterations).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。