用语义引导的多尺度变换器提升动态环境下的视觉定位鲁棒性
Robust Visual Localization via Semantic-Guided Multi-Scale Transformer
- 通过层级变换器融合多尺度几何与上下文信息
- 在TartanAir数据集上显著优于现有方法,抗动态物体和光照变化
- 适合需要高鲁棒性的自动驾驶与机器人定位场景
视觉定位在动态环境中仍具挑战性,光照波动、恶劣天气和移动物体干扰外观线索。尽管特征表示有所进展,现有绝对位姿回归方法在不同条件下难以保持一致性。为此,我们提出一种结合多尺度特征学习与语义场景理解的框架。该方法采用带跨尺度注意力的层级变换器,融合几何细节与上下文线索,既保留空间精度又适应环境变化。训练时通过神经场景表示引入语义监督,引导网络学习视图不变特征,编码持久结构信息并抑制复杂环境干扰。在TartanAir数据集上的实验表明,该方法在存在动态物体、光照变化和遮挡的挑战场景中显著优于现有位姿回归方法。研究结果表明,将多尺度处理与语义引导结合是实现真实动态环境中鲁棒视觉定位的可行策略。
原文摘要 · Abstract (English)
Visual localization remains challenging in dynamic environments where fluctuating lighting, adverse weather, and moving objects disrupt appearance cues. Despite advances in feature representation, current absolute pose regression methods struggle to maintain consistency under varying conditions. To address this challenge, we propose a framework that synergistically combines multi-scale feature learning with semantic scene understanding. Our approach employs a hierarchical Transformer with cross-scale attention to fuse geometric details and contextual cues, preserving spatial precision while adapting to environmental changes. We improve the performance of this architecture with semantic supervision via neural scene representation during training, guiding the network to learn view-invariant features that encode persistent structural information while suppressing complex environmental interference. Experiments on TartanAir demonstrate that our approach outperforms existing pose regression methods in challenging scenarios with dynamic objects, illumination changes, and occlusions. Our findings show that integrating multi-scale processing with semantic guidance offers a promising strategy for robust visual localization in real-world dynamic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。