arXiv:2512.04303cs.CVcs.AI2025-12中稿 · 3DV 2026被引 1

用单目摄像头精准还原道路微结构,自监督学习无需标注数据。

Gamma-from-Mono: Road-Relative, Metric, Self-Supervised Monocular Geometry for Vehicular Applications

论文配图:Gamma-from-Mono: Road-Relative, Metric, Self-Supervised Monocular Geometry for Vehicular Applications
图 1 · 摘自论文原文
  • 分离全局平面与局部起伏,用伽马值量化高度偏差
  • 在KITTI和RSRD上近场深度误差低于现有方法,参数仅888万
  • 适合车载系统,无需外参标定,物理意义明确

车辆对三维环境的精确感知,尤其是路面凹凸、坡度等细粒度几何特征,对安全舒适控制至关重要。传统单目深度估计常过度平滑这些细节,影响运动规划与稳定性。为此,我们提出轻量级单目几何估计方法Gamma-from-Mono(GfM),通过解耦全局与局部结构,解决单相机重建的投影歧义。GfM预测主导道路平面及残差变化,以伽马(gamma)表示——即某点高于平面的高度与其距相机深度之比,基于经典平面视差几何定义。仅需相机离地高度,该表示可闭式求解度量深度,避免完整外参标定,自然强化近路细节。其物理可解释性使其适合自监督学习,无需大规模标注数据。在KITTI和道路表面重建数据集(RSRD)上,GfM在近场深度与伽马估计上均达当前最优,全局深度表现亦具竞争力。所提模型仅8.88M参数,适应多种相机配置,据我们所知,是首个在RSRD上评估的自监督单目方法。

原文摘要 · Abstract (English)

Accurate perception of the vehicle's 3D surroundings, including fine-scale road geometry, such as bumps, slopes, and surface irregularities, is essential for safe and comfortable vehicle control. However, conventional monocular depth estimation often oversmooths these features, losing critical information for motion planning and stability. To address this, we introduce Gamma-from-Mono (GfM), a lightweight monocular geometry estimation method that resolves the projective ambiguity in single-camera reconstruction by decoupling global and local structure. GfM predicts a dominant road surface plane together with residual variations expressed by gamma, a dimensionless measure of vertical deviation from the plane, defined as the ratio of a point's height above it to its depth from the camera, and grounded in established planar parallax geometry. With only the camera's height above ground, this representation deterministically recovers metric depth via a closed form, avoiding full extrinsic calibration and naturally prioritizing near-road detail. Its physically interpretable formulation makes it well suited for self-supervised learning, eliminating the need for large annotated datasets. Evaluated on KITTI and the Road Surface Reconstruction Dataset (RSRD), GfM achieves state-of-the-art near-field accuracy in both depth and gamma estimation while maintaining competitive global depth performance. Our lightweight 8.88M-parameter model adapts robustly across diverse camera setups and, to our knowledge, is the first self-supervised monocular approach evaluated on RSRD.

单目几何自监督车载感知道路重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。