arXiv:2604.24328cs.CV2026-04

用代数结构提升单目深度估计,让模型更懂透视投影规律。

Monocular Depth Estimation via Neural Network with Learnable Algebraic Group and Ring Structures

论文配图:Monocular Depth Estimation via Neural Network with Learnable Algebraic Group and Ring Structures
图 1 · 摘自论文原文
  • 将特征图建模为流形上的层截面,引入可学习的群、环和层结构。
  • 在KITTI、NYU-Depth V2等数据集上显著优于现有方法,零样本泛化能力强。
  • 适合对几何先验敏感的计算机视觉任务,如自动驾驶、3D重建。

单目深度估计(MDE)近年来得益于卷积神经网络和基于Transformer的架构取得了显著进展。然而,这些方法通常将问题视为欧氏网格上的通用图像到图像回归,忽略了透视投影带来的内在代数与几何结构。为解决这一局限,我们提出LAGRNet,一种从根本上基于代数几何的新型框架,通过在深度学习流程中显式嵌入可学习的群、环和层结构。将特征图建模为近似图像流形上的层截面,该方法首先建立由可学习代数群作用参数化的群定义特征流形(GFM),以实现投影等变性和对视角变化的鲁棒性。为促进代数一致的跨尺度交互,引入环卷积层(RCL),将特征融合形式化为分级环同态。此外,为确保全局拓扑一致性,设计了基于层的模块(SM),通过图像拓扑上的切赫神经网聚合局部深度线索。在KITTI、NYU-Depth V2和ETH3D基准上的广泛零样本评估表明,LAGRNet在准确率和泛化能力上均显著优于当前最优方法。

原文摘要 · Abstract (English)

Monocular depth estimation (MDE) has witnessed remarkable progress driven by Convolutional Neural Networks and transformer-based architectures. However, these approaches typically treat the problem as a generic image-to-image regression on Euclidean grids, thereby overlooking the intrinsic algebraic and geometric structures induced by perspective projection. To address this limitation, we propose LAGRNet, a novel framework that fundamentally grounds MDE in algebraic geometry by explicitly embedding learnable group, ring, and sheaf structures into the deep learning pipeline. Modeling feature maps as sections of a sheaf over an approximated image manifold, our method first establishes a Group-defined Feature Manifold (GFM) parameterized by a learned algebraic group action to enforce projective equivariance and robustness against view changes. To facilitate algebraically consistent cross-scale interactions, we subsequently introduce a Ring Convolution Layer (RCL) that formulates feature fusion as a graded ring homomorphism. Furthermore, to ensure global topological consistency, a Sheaf-based Module (SM) aggregates local depth cues via Čech nerve on the image topology. Extensive zero-shot evaluations across the KITTI, NYU-Depth V2, and ETH3D benchmarks demonstrate that LAGRNet significantly outperforms state-of-the-art methods in both accuracy and generalization capabilities.

深度估计代数几何神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。