通过增强图像外观多样性提升视觉定位鲁棒性
Adversarial Exploitation of Data Diversity Improves Visual Localization
- 用3D高斯点云合成不同光照天气下的多样化训练数据
- 双分支对抗训练使室内误差降低50%、室外降低44%
- 特别适合动态驾驶、昼夜切换等复杂场景应用
视觉定位需估计相机在已知场景中的位姿,是自主系统的基础能力。尽管绝对位姿回归(APR)方法推理高效,但泛化能力不足。现有方法虽通过多视角数据增强改善性能,却忽视了外观多样性的重要性。本文发现外观变化是实现鲁棒定位的关键。我们首先将真实2D图像转换为具有不同外观和去模糊能力的3D高斯点云,从而生成不仅在位姿、更在光照、天气等环境条件上多样化的训练数据。为充分挖掘该数据潜力,构建双分支联合训练框架,并引入对抗判别器以弥合仿真到真实的差距。大量实验表明,本方法显著优于现有最先进方法:在室内数据集上,平移误差降低50%,旋转误差降低41%;在室外数据集上,分别降低38%和44%。尤其在动态驾驶、多变天气及昼夜交替场景中,表现远超以往APR方法。
原文摘要 · Abstract (English)
Visual localization, which estimates a camera's pose within a known scene, is a fundamental capability for autonomous systems. While absolute pose regression (APR) methods have shown promise for efficient inference, they often struggle with generalization. Recent approaches attempt to address this through data augmentation with varied viewpoints, yet they overlook a critical factor: appearance diversity. In this work, we identify appearance variation as the key to robust localization. Specifically, we first lift real 2D images into 3D Gaussian Splats with varying appearance and deblurring ability, enabling the synthesis of diverse training data that varies not just in poses but also in environmental conditions such as lighting and weather. To fully unleash the potential of the appearance-diverse data, we build a two-branch joint training pipeline with an adversarial discriminator to bridge the syn-to-real gap. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods, reducing translation and rotation errors by 50\% and 41\% on indoor datasets, and 38\% and 44\% on outdoor datasets. Most notably, our method shows remarkable robustness in dynamic driving scenarios under varying weather conditions and in day-to-night scenarios, where previous APR methods fail. Project Page: https://ai4ce.github.io/RAP/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。