让CLIP模型学会理解热成像,解决光照不足时的视觉盲区
T-CLIP: Enabling Thermal Perception for Contrastive Language-Image Pretraining

- 构建物理感知的热成像标注流水线与数据集IR-Cap
- 提出解耦双LoRA框架,分别处理场景与物体级热特征
- 在三个热成像基准上均超越现有方法,可拓展至热图生成
热成像在低光和恶劣天气下具有显著优势,但以CLIP为代表的基础视觉-语言模型因存在热感知鸿沟而难以将热图像与文本描述对齐。我们识别出三大挑战:缺乏带注释的热成像数据集、标准大语言模型无法推理热现象、以及热成像中全局场景上下文与物体级热信号在单一嵌入空间中冲突的表征难题。为此,我们提出IR-Cap——首个物理感知的热成像标注流水线与数据集,提供跨三个公开基准的全局与细粒度热描述;并设计T-CLIP,一种解耦双LoRA框架,独立适配CLIP以实现场景级与物体级热理解。T-CLIP在三个热成像基准上的跨模态检索任务中持续优于所有基线,并初步验证其在文本条件热图像生成中的适用性。
原文摘要 · Abstract (English)
Thermal imaging offers a powerful alternative to visible-spectrum vision under challenging conditions such as low illumination and adverse weather, yet foundational vision-language models like CLIP fail to align thermal images with textual descriptions due to a fundamental thermal perception gap. We identify three major challenges: the lack of captioned thermal datasets, the inability of standard LLMs to reason about thermal phenomena, and a key representational challenge in thermal imaging where global scene context and object-level heat signatures conflict when learned together in a single embedding space. To address these, we introduce IR-Cap, the first physics-aware thermal captioning pipeline and dataset providing complementary global and fine-grained thermal descriptions across three public benchmarks, and T-CLIP, a decoupled dual-LoRA framework that independently adapts CLIP for scene-level and object-level thermal understanding. T-CLIP achieves consistent improvements over all baselines across three thermal benchmarks in cross-modal retrieval, and we provide an exploratory demonstration of its applicability to text-conditioned thermal image generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。