用双失焦与立体视觉联合估计深度,实现小体积高精度测距。
Depth from Dual Differential Defocus and Stereo Consensus

- 基于双失焦理论与立体匹配的耦合算法,统一深度估计
- 4毫米基线下实现0.3-1.64米范围、1厘米误差的深度图
- 适合小型化被动双目测距设备,如移动机器人或穿戴设备
我们提出D^3S Consensus,一种基于物理模型的闭式算法,融合深度从失焦(DfD)与立体视觉,在超出相机景深范围的长工作距离内实现高精度深度估计。给定一对双失焦立体图像,该方法利用新型双差分失焦(D^3)理论与立体信息,联合估计过约束深度集,并通过物理独立线索间的共识机制筛选最可信结果,剔除不可靠估计。分析表明,在相同误差容忍度下,该方法仅需传统三角测量系统1/10的基线。我们实现了首个原型系统,基线仅4毫米,等效焦距12毫米,可生成高达900×1800像素的深度图,在0.3至1.64米范围内平均绝对误差为1厘米,单次拍摄完成。其精度已超越部分商用大体积立体相机。
原文摘要 · Abstract (English)
We introduce D^3S Consensus, a physics-based, closed-form algorithm that unifies depth-from-defocus (DfD) and stereo to achieve highly accurate depth estimation throughout an extended working range beyond the depth-of-field (DoF) of cameras. Given a pair of dual-defocus stereo images, the method estimates an overdetermined set of depth using a novel DfD theory, Dual Differential Defocus (D^3), and (S)tereo in a coupled fashion. It then picks the most confident depth prediction from the set by enforcing consensus between these physically independent cues to reject unreliable estimates. Analysis shows that D^3S achieves a comparable working range under the same error tolerance with 10x smaller baseline than previous triangulation-based depth estimation systems. This enables compact passive binocular rangefinders with substantially smaller form factors than conventional stereo and DfD designs. We demonstrate the first D^3S prototype with only 4 mm baseline and 12 mm EFL. It generates up to 900 x 1800-pixel depth maps with 1-cm mean absolute error over 0.3-1.64 m from a snapshot acquisition. This has surpassed the reported accuracy of certain commercially available stereo cameras with much larger form factors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。