用视觉变压器提升热成像重定位精度,适应大场景复杂环境。
ThermalLoc: A Vision Transformer-Based Approach for Robust Thermal Camera Relocalization in Large-Scale Environments
- 结合EfficientNet与Transformer提取热图像局部和全局特征
- 在两个数据集上优于AtLoc、MapNet等现有方法,精度显著提升
- 适合需要高鲁棒性的大规模热成像定位任务
热成像通过物体热辐射捕捉环境信息,与依赖针孔成像的可见光相机机制本质不同,导致传统视觉重定位方法难以直接应用于热图像。尽管深度学习在相机重定位领域取得进展,但专为热成像设计的方法仍较少。为此,我们提出ThermalLoc,一种基于EfficientNet与Transformer的端到端热图像重定位方法,通过双MLP网络实现绝对位姿回归。在公开热惯性里程计数据集及自建数据集上的评估表明,ThermalLoc优于AtLoc、MapNet、PoseNet和RobustLoc等代表性方法,在准确性和鲁棒性方面均表现更优。
原文摘要 · Abstract (English)
Thermal cameras capture environmental data through heat emission, a fundamentally different mechanism compared to visible light cameras, which rely on pinhole imaging. As a result, traditional visual relocalization methods designed for visible light images are not directly applicable to thermal images. Despite significant advancements in deep learning for camera relocalization, approaches specifically tailored for thermal camera-based relocalization remain underexplored. To address this gap, we introduce ThermalLoc, a novel end-to-end deep learning method for thermal image relocalization. ThermalLoc effectively extracts both local and global features from thermal images by integrating EfficientNet with Transformers, and performs absolute pose regression using two MLP networks. We evaluated ThermalLoc on both the publicly available thermal-odometry dataset and our own dataset. The results demonstrate that ThermalLoc outperforms existing representative methods employed for thermal camera relocalization, including AtLoc, MapNet, PoseNet, and RobustLoc, achieving superior accuracy and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。