arXiv:2607.06552cs.CV2026-07被引 1

构建首个红外遥感视觉语言数据集,提升模型对热成像特征的理解。

MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation

论文配图:MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation
图 1 · 摘自论文原文
  • 基于CLIP风格对比学习与VLM指令微调,专为红外图像设计
  • 合成60万张红外图,使模型在热成像理解上比灰度转换提升显著
  • 适合遥感、多模态学习研究者,推动红外视觉语言对齐

红外遥感影像能捕捉强度结构、目标-背景对比及光照不变特征,这些在可见光图像中常不可见。然而,现有遥感视觉语言资源和模型多聚焦于可见光语义,导致红外视觉语言理解研究不足。本文提出MonoIR-RS,一个大规模红外遥感视觉语言数据集与基准,结合红外感知的数据构建、CLIP式对比适应及VLM指令微调。该数据集源自同一数据源,与FusionRS同分割,保留红外图像作为模型输入模态,生成60万张合成红外图像和59,032条保留红外语义的文本标注。实验使用其中带语言监督的子集,其标注围绕灰度结构与红外对比特征重写,而非RGB外观。在AVIID基准上,合成红外图像相比灰度转换更接近真实热成像。我们微调了五种CLIP骨干与六种VLM骨干,并与零样本表现对比:红外感知适应使CLIP平均召回率提升12.8点,同时使VLM对红外线索的覆盖率达100%,残余RGB颜色泄漏趋近于零。通过将红外模态从RGB-IR双模态学习中分离,MonoIR-RS提供了一个可控、可复现的测试平台,用于对齐红外遥感证据与语言。

原文摘要 · Abstract (English)

Infrared remote-sensing imagery captures intensity structure, object-background contrast, and illumination-invariant cues often invisible in RGB imagery. Yet, most remote-sensing vision-language resources and models focus on visible-band semantics, leaving infrared vision-language understanding underexplored. We introduce MonoIR-RS, a large-scale infrared remote-sensing vision-language dataset and benchmark that couples IR-aware data construction with CLIP-style contrastive adaptation and VLM instruction tuning. Built from the same source pool and split as FusionRS, MonoIR-RS retains the infrared image as the model-facing modality, yielding 600,000 synthesized infrared images and 59,032 retained IR-aware caption records. The model experiments use this retained language-supervision subset, whose captions rewrite supervision around grayscale structure and infrared-style contrast instead of RGB appearance. We show that the synthesized infrared imagery is markedly closer to real thermal imagery than a grayscale conversion on the AVIID benchmark. We fine-tune five CLIP backbones and six VLM backbones, and calibrate them against zero-shot behavior: IR-aware adaptation lifts CLIP mean recall by up to 12.8 points and drives VLM captioning IR-cue coverage to 100% while reducing residual RGB-color leakage to near zero. By isolating the infrared modality from RGB-IR dual-modal learning, MonoIR-RS offers a controlled, reproducible testbed for aligning infrared remote-sensing evidence with language.

红外遥感视觉语言多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。