用轻量框架修复单目深度估计在透明场景中的误差,提升机器人感知可靠性。
OptiGeo: Efficient Monocular Geometry for Embodied Perception in Optically Challenging Scenes

- 通过教师模型纠正传感器偏差,实现对透明区域的清洁几何监督
- 仅30M参数即超越300M大模型和数十亿参数多视角基线
- 适合需要高效部署的机器人导航系统,尤其在光学复杂场景
单目深度估计虽具备强开放域泛化能力,但在透明、反光和镜面环境中仍难以可靠部署,因深度传感器常产生缺失或偏差。现有方法依赖场景特定预处理、辅助模块或后期微调,增加冗余且易过拟合。本文将此问题视为基础模型训练中的局部失效模式,识别出传感器诱导的监督偏差是关键瓶颈:模型从有缺陷的真实深度监督中继承了光学复杂区域的错误模式。为此提出OptiGeo,一种感知偏差的训练框架,利用清洁几何教师模型与残差裁剪对齐,重构透明目标的几何结构。将透明场景渲染作为紧凑的清洁几何来源,而非大规模领域微调数据集。仅需少量目标渲染数据,即可修正真实传感器无法可靠监督的局部几何失真。尽管参数仅30M,OptiGeo在透明场景基准上显著优于300M级单目模型及百亿级多视角基线,同时保持零样本深度与边界锐度的竞争力。真实导航实验进一步验证其在光学挑战场景下的实用性。
原文摘要 · Abstract (English)
Monocular depth estimation has achieved strong open-domain generalization, yet reliable robotic deployment remains difficult in transparent, reflective, and specular environments, where depth sensors often produce missing or biased depth. Existing methods often handle such optical failures with scene-specific preprocessing, auxiliary modules, or post-hoc fine-tuning. While effective in constrained settings, these designs increase architectural redundancy and can over-specialize general geometry models to narrow optical scenarios. We revisit this problem as a localized failure mode within base-model training and identify sensor-induced supervision bias as a key bottleneck: models inherit sensor failure patterns from biased real-depth supervision in optically challenging regions. We then introduce OptiGeo, a bias-aware training framework that rehabilitates biased real supervision using a clean-geometry teacher and residual-trimmed alignment. We redefine transparency-targeted rendering as a compact source of clean optical geometry, rather than a large domain-specific fine-tuning set. With only a small targeted rendering set, OptiGeo learns the geometric structure of transparent objects and regions, correcting local geometry distortions that real sensors cannot reliably supervise. Despite only 30M parameters, OptiGeo outperforms substantially larger 300M-scale monocular models and billion-scale multi-view baselines on transparent-scene benchmarks, while remaining competitive on general zero-shot depth and boundary sharpness. Real-world navigation cases further validate its practicality as an efficient perception module in optically challenging scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。