arXiv:2507.14449cs.CV2025-07ICCV被引 14

首个面向真实红外图像的多模态大模型,解决数据稀缺与模态差异问题。

IRGPT: Understanding Real-world Infrared Image with Bi-cross-modal Curriculum on Large-scale Benchmark

  • 基于26万对真实红外-文本配对数据,构建首个大规模红外图文数据集。
  • 提出双向跨模态课程学习策略,显著提升红外图像理解能力。
  • 在9项任务上超越更大模型,适合红外视觉与多模态研究者使用。

真实世界红外图像因缺乏对齐文本数据和特定域特征,给视觉语言模型带来独特挑战。现有方法依赖通过可见光图像风格迁移生成的合成红外图像,难以捕捉红外模态的独特性。为此,我们提出IRGPT,首个面向真实红外图像的多模态大语言模型,基于包含超过26万对真实红外-文本配对的大规模红外-文本数据集(IR-TD)。该数据集中的文本由两种互补方式生成:(1)大语言模型对可见光图像的生成描述;(2)基于规则的标注描述。此外,我们引入一种双向跨模态课程迁移学习策略,系统地从可见光域向红外域迁移知识,同时考虑红外-可见与红外-文本的难度评分。在涵盖9个任务(如识别、定位)的基准测试中,IRGPT表现达到当前最优,甚至优于更大规模模型。

原文摘要 · Abstract (English)

Real-world infrared imagery presents unique challenges for vision-language models due to the scarcity of aligned text data and domain-specific characteristics. Although existing methods have advanced the field, their reliance on synthetic infrared images generated through style transfer from visible images, which limits their ability to capture the unique characteristics of the infrared modality. To address this, we propose IRGPT, the first multi-modal large language model for real-world infrared images, built upon a large-scale InfraRed-Text Dataset (IR-TD) comprising over 260K authentic image-text pairs. The proposed IR-TD dataset contains real infrared images paired with meticulously handcrafted texts, where the initial drafts originated from two complementary processes: (1) LLM-generated descriptions of visible images, and (2) rule-based descriptions of annotations. Furthermore, we introduce a bi-cross-modal curriculum transfer learning strategy that systematically transfers knowledge from visible to infrared domains by considering the difficulty scores of both infrared-visible and infrared-text. Evaluated on a benchmark of 9 tasks (e.g., recognition, grounding), IRGPT achieves state-of-the-art performance even compared with larger-scale models.

红外图像多模态大模型视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。