用两张图实现高精度深度重建,突破传统SFF对大量图像的依赖。
Towards Minimal Focal Stack in Shape from Focus
- 通过物理模型生成全焦图和能量差图,扩充仅两张输入的焦点堆栈。
- 在合成与真实数据集上达到现有顶尖方法的精度水平。
- 适合资源受限场景下的高效三维重建,如移动设备或嵌入式系统。
基于聚焦(SFF)是一种通过分析焦点堆栈中图像的聚焦变化来估计场景结构的深度重建技术,焦点堆栈是不同对焦设置下拍摄的一系列图像。现有SFF方法的关键局限在于对密集采样、大尺寸焦点堆栈的依赖,限制了其实际应用。本文提出一种焦点堆栈增强方法,使SFF模型仅用两张图像即可实现深度估计且不损失精度。我们引入一种简单而有效的物理驱动增强策略,生成两个辅助线索:由两张输入图像估计出的全焦图(AiF),以及计算为全焦图与输入图像差异能量的能量差图(EOD)。同时,提出一个深度网络,从增强后的焦点堆栈中计算深度体,并利用多尺度卷积门控循环单元(ConvGRUs)进行迭代精炼。在合成与真实世界数据集上的大量实验表明,该增强方法可显著提升现有最先进SFF模型的表现,使其在极小堆栈规模下仍保持领先性能。
原文摘要 · Abstract (English)
Shape from Focus (SFF) is a depth reconstruction technique that estimates scene structure from focus variations observed across a focal stack, that is, a sequence of images captured at different focus settings. A key limitation of SFF methods is their reliance on densely sampled, large focal stacks, which limits their practical applicability. In this study, we propose a focal stack augmentation that enables SFF methods to estimate depth using a reduced stack of just two images, without sacrificing precision. We introduce a simple yet effective physics-based focal stack augmentation that enriches the stack with two auxiliary cues: an all-in-focus (AiF) image estimated from two input images, and Energy-of-Difference (EOD) maps, computed as the energy of differences between the AiF and input images. Furthermore, we propose a deep network that computes a deep focus volume from the augmented focal stacks and iteratively refines depth using convolutional Gated Recurrent Units (ConvGRUs) at multiple scales. Extensive experiments on both synthetic and real-world datasets demonstrate that the proposed augmentation benefits existing state-of-the-art SFF models, enabling them to achieve comparable accuracy. The results also show that our approach maintains state-of-the-art performance with a minimal stack size.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。