arXiv:2512.10498cs.CV2025-12IJCV被引 5

用多尺度方向膨胀拉普拉斯增强聚焦信息,提升深度估计精度与鲁棒性。

Robust Shape from Focus via Multiscale Directional Dilated Laplacian and Recurrent Network

  • 采用手工设计的多尺度方向膨胀拉普拉斯核提取鲁棒聚焦特征
  • 通过轻量级循环网络实现逐层精炼,生成高分辨率无伪影深度图
  • 适合需要高精度深度重建的工业检测与三维成像场景

基于聚焦的形状(SFF)是一种被动深度估计技术,通过分析焦堆栈中的聚焦变化来推断场景深度。近期大多数基于深度学习的SFF方法通常分为两阶段:首先使用复杂的特征编码器提取聚焦体积(即每个像素在焦堆栈中聚焦可能性的表示);随后采用简单的单步聚合技术估计深度,常引入伪影并放大深度图中的噪声。为此,我们提出一种混合框架:利用手工设计的多尺度方向膨胀拉普拉斯(DDL)核传统计算多尺度聚焦体积,捕捉长程和方向性聚焦变化,形成更鲁棒的聚焦表示;这些体积输入到一个轻量级、多尺度的GRU-based深度提取模块,该模块在低分辨率下迭代优化初始深度估计,提升计算效率;最后,嵌入在循环网络中的可学习凸上采样模块重建高分辨率深度图,同时保留精细场景细节与锐利边界。在合成与真实世界数据集上的大量实验表明,本方法优于现有先进深度学习与传统方法,在多种焦距条件下均实现更高精度与更强泛化能力。

原文摘要 · Abstract (English)

Shape-from-Focus (SFF) is a passive depth estimation technique that infers scene depth by analyzing focus variations in a focal stack. Most recent deep learning-based SFF methods typically operate in two stages: first, they extract focus volumes (a per pixel representation of focus likelihood across the focal stack) using heavy feature encoders; then, they estimate depth via a simple one-step aggregation technique that often introduces artifacts and amplifies noise in the depth map. To address these issues, we propose a hybrid framework. Our method computes multi-scale focus volumes traditionally using handcrafted Directional Dilated Laplacian (DDL) kernels, which capture long-range and directional focus variations to form robust focus volumes. These focus volumes are then fed into a lightweight, multi-scale GRU-based depth extraction module that iteratively refines an initial depth estimate at a lower resolution for computational efficiency. Finally, a learned convex upsampling module within our recurrent network reconstructs high-resolution depth maps while preserving fine scene details and sharp boundaries. Extensive experiments on both synthetic and real-world datasets demonstrate that our approach outperforms state-of-the-art deep learning and traditional methods, achieving superior accuracy and generalization across diverse focal conditions.

深度估计聚焦感知图像处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。