arXiv:2604.06124cs.CVcs.AI2026-04

轻量级适配让预训练视觉模型学会识别热成像中的物种和栖息地特征。

Lightweight Multimodal Adaptation of Vision Language Models for Species Recognition and Habitat Context Interpretation in Drone Thermal Imagery

  • 用轻量投影器对齐可见光与热成像特征,实现跨模态迁移。
  • 大象识别F1达0.968,犀牛计数准确率100%,表现最优。
  • 可同时分析物种与栖息地信息,适合生态监测场景。

本研究提出一种轻量级多模态适配框架,弥合了基于可见光预训练的视觉语言模型(VLMs)与热红外图像之间的表征差距,并在真实无人机采集的数据集上验证其实用性。通过无人机采集影像构建热成像数据集,利用多模态投影器对齐进行微调,使模型能将可见光视觉表征迁移到热辐射输入中。在封闭集与开放集提示条件下,对InternVL3-8B-Instruct、Qwen2.5-VL-7B-Instruct和Qwen3-VL-8B-Instruct三款代表性模型进行了物种识别与实例计数的基准测试。其中,Qwen3-VL-8B-Instruct在开放集提示下表现最佳,对鹿、犀牛、大象的F1分数分别为0.935、0.915、0.968,计数准确率(within-1)分别为0.779、0.982、1.000。此外,结合同步获取的可见光与热成像,模型可生成包括地表覆盖特征、关键地貌要素及人类干扰痕迹在内的栖息地上下文信息。结果表明,基于轻量投影器的适配方法为将预训练视觉语言模型应用于无人机热成像提供了高效可行路径,推动其从目标识别扩展至生态监测中的栖息地语境理解。

原文摘要 · Abstract (English)

This study proposes a lightweight multimodal adaptation framework to bridge the representation gap between RGB-pretrained VLMs and thermal infrared imagery, and demonstrates its practical utility using a real drone-collected dataset. A thermal dataset was developed from drone-collected imagery and was used to fine-tune VLMs through multimodal projector alignment, enabling the transfer of information from RGB-based visual representations to thermal radiometric inputs. Three representative models, including InternVL3-8B-Instruct, Qwen2.5-VL-7B-Instruct, and Qwen3-VL-8B-Instruct, were benchmarked under both closed-set and open-set prompting conditions for species recognition and instance enumeration. Among the tested models, Qwen3-VL-8B-Instruct with open-set prompting achieved the best overall performance, with F1 scores of 0.935 for deer, 0.915 for rhino, and 0.968 for elephant, and within-1 enumeration accuracies of 0.779, 0.982, and 1.000, respectively. In addition, combining thermal imagery with simultaneously collected RGB imagery enabled the model to generate habitat-context information, including land-cover characteristics, key landscape features, and visible human disturbance. Overall, the findings demonstrate that lightweight projector-based adaptation provides an effective and practical route for transferring RGB-pretrained VLMs to thermal drone imagery, expanding their utility from object-level recognition to habitat-context interpretation in ecological monitoring.

多模态热成像生态监测轻量适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。