arXiv:2604.06576cs.CVeess.IV2026-04中稿 · IEEE Transactions …被引 2

用数学框架提升单目深度估计精度,尤其改善边缘处的预测效果。

LiftFormer: Lifting and Frame Theory Based Monocular Depth Estimation Using Depth and Edge Oriented Subspace Representation

论文配图:LiftFormer: Lifting and Frame Theory Based Monocular Depth Estimation Using Depth and Edge Oriented Subspace Representation
图 1 · 摘自论文原文
  • 基于提升理论构建深度导向子空间,连接图像颜色与深度值
  • 在边缘区域引入感知子空间,显著降低深度突变处的误差
  • 在多个基准数据集上达到顶尖性能,适合3D视觉任务研究者

单目深度估计(MDE)近年来受到广泛关注,因其在三维视觉中的关键作用。MDE是从单张图像或视频中估计深度图以表示场景的三维结构,这是一个高度病态的问题。本文提出一种基于提升理论拓扑的LiftFormer,用于构建一个中间子空间,连接图像颜色特征与深度值,并构建一个增强边缘附近深度预测的子空间。将深度值预测问题转化为深度导向几何表示(DGR)子空间特征表示,从而实现从颜色值到几何深度值的学习桥梁。利用深度分箱中的线性相关向量,基于框架理论构建冗余且鲁棒的DGR子空间,将图像空间特征转换至该子空间,使其直接对应深度值。此外,考虑到边缘通常表现为深度图中的剧烈变化,且易被误预测,本文还构建了边缘感知表示(ER)子空间,将深度特征在此处进行转换并进一步增强边缘附近的局部特征。实验结果表明,LiftFormer在多个常用数据集上达到最先进性能,消融实验证明了所提提升模块的有效性。

原文摘要 · Abstract (English)

Monocular depth estimation (MDE) has attracted increasing interest in the past few years, owing to its important role in 3D vision. MDE is the estimation of a depth map from a monocular image/video to represent the 3D structure of a scene, which is a highly ill-posed problem. To solve this problem, in this paper, we propose a LiftFormer based on lifting theory topology, for constructing an intermediate subspace that bridges the image color features and depth values, and a subspace that enhances the depth prediction around edges. MDE is formulated by transforming the depth value prediction problem into depth-oriented geometric representation (DGR) subspace feature representation, thus bridging the learning from color values to geometric depth values. A DGR subspace is constructed based on frame theory by using linearly dependent vectors in accordance with depth bins to provide a redundant and robust representation. The image spatial features are transformed into the DGR subspace, where these features correspond directly to the depth values. Moreover, considering that edges usually present sharp changes in a depth map and tend to be erroneously predicted, an edge-aware representation (ER) subspace is constructed, where depth features are transformed and further used to enhance the local features around edges. The experimental results demonstrate that our LiftFormer achieves state-of-the-art performance on widely used datasets, and an ablation study validates the effectiveness of both proposed lifting modules in our LiftFormer.

深度估计提升理论边缘感知单目视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。