无需微调,一套模型适配多种热成像任务与环境。
AnyThermal: Towards Learning Universal Representations for Thermal Perception
- 用多源热成像数据蒸馏视觉大模型特征,构建通用热感知编码器。
- 在4类场景下跨任务测试,性能提升最高达36%。
- 开源同步采集平台与数据集,推动热成像研究标准化。
我们提出 AnyThermal,一种可捕捉鲁棒、任务无关热成像特征的主干网络,适用于跨模态定位、热图像分割和单目深度估计等多种任务。现有热成像主干网络依赖小规模数据进行特定任务训练,适用范围受限。不同于以往方法,AnyThermal可在室内、航拍、非铺装路面和城市等多样环境中,无需任务微调即可通用。核心思路是利用来自多环境的热成像数据,将视觉基础模型(如 DINOv2)的特征知识蒸馏至热成像编码器。为弥合现有 RGB-Thermal 数据集间的差异,我们推出首个开源同步采集平台 TartanRGBT,用于收集跨场景的同步 RGB-Thermal 图像。基于此平台构建了 TartanRGBT 数据集——一个覆盖4种环境、平衡且多样化的数据集。实验表明,AnyThermal 与 TartanRGBT 在多个现有数据集上实现了最先进的性能,各项任务平均提升最高达36%。
原文摘要 · Abstract (English)
We present AnyThermal, a thermal backbone that captures robust task-agnostic thermal features suitable for a variety of tasks such as cross-modal place recognition, thermal segmentation, and monocular depth estimation using thermal images. Existing thermal backbones that follow task-specific training from small-scale data result in utility limited to a specific environment and task. Unlike prior methods, AnyThermal can be used for a wide range of environments (indoor, aerial, off-road, urban) and tasks, all without task-specific training. Our key insight is to distill the feature representations from visual foundation models such as DINOv2 into a thermal encoder using thermal data from these multiple environments. To bridge the diversity gap of the existing RGB-Thermal datasets, we introduce the TartanRGBT platform, the first open-source data collection platform with synced RGB-Thermal image acquisition. We use this payload to collect the TartanRGBT dataset - a diverse and balanced dataset collected in 4 environments. We demonstrate the efficacy of AnyThermal and TartanRGBT, achieving state-of-the-art results with improvements of up to 36% across diverse environments and downstream tasks on existing datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。