提出高效合成模糊图像方法,解决真实相机成像难题。
Efficient Depth- and Spatially-Varying Image Simulation for Defocus Deblur
- 联合建模深度相关模糊与空间变化像差,模拟真实成像过程
- 用低分辨率合成数据训练的模型可泛化至1200万像素真实图像
- 无需真实数据微调,适合智能眼镜等小型设备使用
现代大光圈相机常因景深过浅导致离焦区域图像模糊,这对固定对焦设备(如智能眼镜)尤为不利,因受体积和功耗限制难以加入自动对焦。由于每种相机系统具有独特的光学像差和离焦特性,现有开源数据集训练的深度学习模型在真实场景中普遍存在域偏移问题。本文提出一种高效可扩展的数据集生成方法,无需依赖真实数据微调。该方法同时建模深度相关的模糊效应和空间变化的光学像差,兼顾计算效率与高质量RGB-D数据稀缺的问题。实验表明,基于低分辨率合成图像训练的网络能有效泛化到不同场景下的高分辨率(12MP)真实图像。
原文摘要 · Abstract (English)
Modern cameras with large apertures often suffer from a shallow depth of field, resulting in blurry images of objects outside the focal plane. This limitation is particularly problematic for fixed-focus cameras, such as those used in smart glasses, where adding autofocus mechanisms is challenging due to form factor and power constraints. Due to unmatched optical aberrations and defocus properties unique to each camera system, deep learning models trained on existing open-source datasets often face domain gaps and do not perform well in real-world settings. In this paper, we propose an efficient and scalable dataset synthesis approach that does not rely on fine-tuning with real-world data. Our method simultaneously models depth-dependent defocus and spatially varying optical aberrations, addressing both computational complexity and the scarcity of high-quality RGB-D datasets. Experimental results demonstrate that a network trained on our low resolution synthetic images generalizes effectively to high resolution (12MP) real-world images across diverse scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。