零样本通用视觉里程计,无需标定即可跨场景稳定运行
ZeroVO: Visual Odometry with Minimal Assumptions
- 无标定几何感知网络,自动处理深度与相机参数噪声
- 在KITTI、nuScenes等3个基准上提升超30%性能
- 适合无人车等需快速部署的复杂真实场景
我们提出ZeroVO,一种新型视觉里程计算法,可在不同相机和环境中实现零样本泛化,克服了现有方法依赖预设或固定相机标定的局限。核心创新包括:1)设计无标定、几何感知的网络结构,可处理估计深度和相机参数中的噪声;2)引入基于语言的先验,注入语义信息以增强特征提取并提升对未见场景的泛化能力;3)开发灵活的半监督训练范式,通过无标签数据迭代适应新场景,进一步提升在多样化真实场景中的泛化性能。我们在复杂自动驾驶场景中评估,相比之前方法在KITTI、nuScenes、Argoverse 2三个标准基准上均提升超过30%,并在基于GTA生成的高保真合成数据集上验证有效性。该方法无需微调或相机标定,显著拓展了视觉里程计在大规模真实部署中的适用性。
原文摘要 · Abstract (English)
We introduce ZeroVO, a novel visual odometry (VO) algorithm that achieves zero-shot generalization across diverse cameras and environments, overcoming limitations in existing methods that depend on predefined or static camera calibration setups. Our approach incorporates three main innovations. First, we design a calibration-free, geometry-aware network structure capable of handling noise in estimated depth and camera parameters. Second, we introduce a language-based prior that infuses semantic information to enhance robust feature extraction and generalization to previously unseen domains. Third, we develop a flexible, semi-supervised training paradigm that iteratively adapts to new scenes using unlabeled data, further boosting the models' ability to generalize across diverse real-world scenarios. We analyze complex autonomous driving contexts, demonstrating over 30% improvement against prior methods on three standard benchmarks, KITTI, nuScenes, and Argoverse 2, as well as a newly introduced, high-fidelity synthetic dataset derived from Grand Theft Auto (GTA). By not requiring fine-tuning or camera calibration, our work broadens the applicability of VO, providing a versatile solution for real-world deployment at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。