arXiv:2511.20853cs.CVcs.AI2025-11

首个高分辨率多镜头相机实拍数据集,支持深度模糊与散焦去模糊算法评测。

MODEST: Multi-Optics Depth-of-Field Stereo Dataset

  • 系统性改变焦距与光圈参数,采集真实场景的立体图像
  • 20万张图像覆盖50种光学配置,分辨率高达2000万像素
  • 专为专业相机光学特性设计,适合评估真实世界视觉模型

当前先进计算机视觉算法在可靠浅景深渲染与散焦去模糊方面的训练与评估,受限于缺乏大规模、全帧、高保真、真实图像数据集。浅景深与散焦模糊的光学效应依赖于焦距与光圈等相机配置,需在参数系统变化下严格评估模型表现。此外,现代应用如AR、VR、智能手机、工业机器人等普遍采用立体或多摄像头系统。本文提出MODEST——首个超高清(5472×3648像素,20MP)、多光学配置的浅景深立体DSLR数据集,系统性地在10个复杂真实场景中变换焦距(28-70mm)与光圈(f/22-f/2.8),覆盖多个立体视角。共包含20,000张图像,涵盖50种光学配置,真实还原专业相机系统的光学复杂性。每组场景均包含反射表面、透明玻璃、精细细节、点光源及多尺度深度错觉等挑战元素。同时提供内参与外参标定图像,支持不断演进的标定方法。我们评估了若干前沿浅景深与散焦去模糊方法,并揭示其失败案例与局限性。全面调优分析表明,MODEST能有效评估现有方法对输入图像实际光学配置的敏感性。本工作旨在弥合合成低分辨率训练数据与真实高分辨率相机光学推理泛化之间的现实差距。

原文摘要 · Abstract (English)

Training and evaluation of state-of-the-art computer vision algorithms for reliable shallow depth of field (DoF) rendering and defocus deblurring remain constrained by a persistent lack of large-scale, full-frame, high fidelity, real-image datasets. Optical effects of shallow DoF and defocus blur depend intimately on camera optical configuration set with focal length and aperture; requiring rigorous evaluation of the models when these parameters systematically change. Further, modern applications such as AR, VR, smartphones, industrial robots, etc. deploy stereo or multi-camera systems. We present MODEST - the first ultra-high-resolution(5472x3648px, 20MP), multi-optics depth of field stereo DSLR dataset that methodically varies focal length and aperture for a series of complex, real-world scenes, capturing the optical realism and complexity of professional camera systems. With 20,000 images across 50 distinct optical configurations, focal length in 28-70mm, aperture in f/22-f/2.8 for multiple stereo viewpoints for 10 scenes; this ultra-high-resolution, full-range optics coverage enables controlled analysis of geometric and optical effects for shallow DoF rendering and defocus deblurring. Each scene is curated to have challenging visual elements: reflective surfaces, transparent glass walls, fine-grained details, point lights, and multi-scale depth illusions. In addition, we provide intrinsics and extrinsics calibration images to support ever-evolving calibration methods. We evaluate several SOTA DoF and defocus deblurring methods and demonstrate failure cases and limitations. Our comprehensive tuning analysis demonstrates how MODEST evaluates sensitivity of the SOTA DoF models for the actual focal configuration of input images. This work attempts to bridge the realism gap between synthetic, low-resolution training data and inference generalization on high-resolution real-camera optics.

数据集立体视觉深度模糊真实图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。