用鱼眼与针孔镜头实时估算精确深度,仅需少量参数。
X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

- 通过可学习校准令牌和雅可比畸变偏置,统一不同镜头投影空间。
- 模型仅0.04亿参数,达41帧/秒,深度误差降低25.4%。
- 适合需要多相机实时深度感知的自动驾驶与机器人场景。
我们提出X-Lens,一种紧凑的前馈模型,可从数量可变的校准鱼眼与针孔视角中进行度量深度估计。为支持实时下游感知,模型基于几何感知的异构相机框架,包含两个关键组件:可学习校准令牌实现鱼眼与针孔投影空间的粗对齐;雅可比参数化畸变偏置注入交叉注意力模块,捕捉局部投影变化并促进跨相机一致性,实现仅0.04B参数下高达41 FPS的鲁棒泛化。模型同时预测稠密深度与全局度量尺度,避免增加计算与优化复杂性的辅助重建目标。为实现大规模跨相机泛化与深度学习,模型在多个公开数据集及我们新发布的大型合成数据集OmniScene上训练,该数据集含约266K同步六视角帧、170万张独立图像和103个室内外场景。大量实验在真实与合成的室内外数据集上验证了其优越性能:在OmniScene-Full上,相对绝对误差(AbsRel)较最强基线降低25.4%,参数量减少88.9%,且在传统鱼眼或针孔单一设置下表现具有竞争力。
原文摘要 · Abstract (English)
We present X-lens, a compact feed-forward model for metric depth estimation from a variable number of calibrated fisheye and pinhole views. To support real-time downstream perception, X-lens is built around a geometry-aware heterogeneous camera formulation with two key components. Learnable calibration tokens provide a coarse alignment between fisheye and pinhole projective spaces, while a Jacobian-parameterized distortion bias injected into cross-attention models local projection changes and promotes cross-camera consistency, enabling robust generalization with only 0.04B parameters and up to 41 FPS. The model predicts dense depth together with a global metric scale, avoiding auxiliary reconstruction targets that increase computation and optimization complexity. To learn such cross-camera generalization at scale and depth, X-lens is trained on multiple public datasets and OmniScene, our newly released large-scale synthetic dataset containing approximately 266K synchronized six-view frames, 1.7M individual images, and 103 indoor and outdoor scenes. Extensive experiments on both real-world and synthetic indoor and outdoor datasets demonstrate superior heterogeneous-camera metric depth accuracy, reducing AbsRel by 25.4\% on OmniScene-Full over the strongest baseline while using 88.9\% fewer parameters, with competitive performance on conventional fisheye-only and pinhole-only settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。