arXiv:2512.15971cs.CV2025-12

用视觉语言模型实现少样本多光谱目标检测,提升数据效率

From Words to Wavelengths: VLMs for Few-Shot Multispectral Object Detection

  • 利用VLM融合文本、可见光与热成像三模态信息
  • 在FLIR和M3FD上少样本性能超越专用多光谱模型
  • 适合资源有限但需强鲁棒感知的自动驾驶场景

多光谱目标检测对自动驾驶、监控等安全敏感应用至关重要,需在复杂光照下保持鲁棒感知。然而标注多光谱数据稀缺,制约深度检测器训练。本文借助视觉语言模型(VLM)在计算机视觉中的成功,探索其在少样本多光谱检测中的潜力。我们适配了两种代表性VLM检测器——Grounding DINO与YOLO-World,处理多光谱输入,并提出有效机制融合文本、可见光与热成像模态。在两个主流多光谱图像基准数据集FLIR和M3FD上进行大量实验表明,VLM-based检测器在少样本设置下显著优于使用相近数据训练的专用多光谱模型,且在全监督条件下也达到竞争力或更优性能。研究发现,大规模VLM学习到的语义先验能有效迁移至未见光谱模态,为数据高效多光谱感知提供强大路径。

原文摘要 · Abstract (English)

Multispectral object detection is critical for safety-sensitive applications such as autonomous driving and surveillance, where robust perception under diverse illumination conditions is essential. However, the limited availability of annotated multispectral data severely restricts the training of deep detectors. In such data-scarce scenarios, textual class information can serve as a valuable source of semantic supervision. Motivated by the recent success of Vision-Language Models (VLMs) in computer vision, we explore their potential for few-shot multispectral object detection. Specifically, we adapt two representative VLM-based detectors, Grounding DINO and YOLO-World, to handle multispectral inputs and propose an effective mechanism to integrate text, visual and thermal modalities. Through extensive experiments on two popular multispectral image benchmarks, FLIR and M3FD, we demonstrate that VLM-based detectors not only excel in few-shot regimes, significantly outperforming specialized multispectral models trained with comparable data, but also achieve competitive or superior results under fully supervised settings. Our findings reveal that the semantic priors learned by large-scale VLMs effectively transfer to unseen spectral modalities, ofFering a powerful pathway toward data-efficient multispectral perception.

多光谱检测视觉语言模型少样本学习自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。