arXiv:2608.07015cs.CV2026-08

用语言理解先于检测,让模型跨域识别红外小目标

Understand Before Detect: Vision--Language Learning for Omni-Domain Infrared Small Target Detection

论文配图:Understand Before Detect: Vision--Language Learning for Omni-Domain Infrared Small Target Detection
图 1 · 摘自论文原文
  • 先通过语言监督学习红外目标整体语义,再做精准检测
  • 在6个任务39000+标注上实现跨域泛化,性能显著提升
  • 适合做红外监控、无人机感知等跨场景小目标检测

全领域红外小目标(IRST)检测对红外监视至关重要,但受成像域异质性和目标特征不一致影响,仍具挑战。现有深度学习方法多基于视觉单模态,采用特定任务的监督学习,将全场景观测简化为稀疏目标标注,忽略了跨域不变的语义信息,导致域偏移下性能下降。为此,本文提出“理解先于检测”新范式,将全领域IRST检测重构为以理解驱动的过程:先通过语言监督建立红外目标的整体理解,再迁移跨域表征用于精确检测。基于此,提出JinSight模型,通过语言语义锚定红外表征,使单一模型可泛化于异构红外域。进一步设计低秩空间中的隐式语义交互(LSI),实现全局语义与局部空间特征的高效融合。为弥补多模态全领域IRST基准缺失,构建首个大规模、高多样性视觉-语言数据集OmniIRST-VL,包含超过39,000个标注,覆盖六类互补指令任务,涵盖场景级理解与目标中心推理。

原文摘要 · Abstract (English)

Omni-domain infrared small target (IRST) detection is crucial for infrared surveillance, yet remains challenging due to heterogeneous imaging domains and inconsistent target characteristics. Previous deep learning-based methods have been developed for visual-only paradigms and achieved promising performance on domain-specific tasks. However, existing methods follow the task-specific supervised learning paradigm. This paradigm simplifies the full-scene infrared observations to sparse target supervision, discarding the semantics that remain invariant across heterogeneous domains. Consequently, detection performance suffers substantially under domain shifts. To handle this issue, we introduce \textbf{``understand before detect''}, a paradigm that formulates omni-domain IRST detection as an understanding-driven process, where holistic infrared target understanding precedes precise detection. Building on this paradigm, we propose \textbf{JinSight}, which first develops holistic IRST understanding through language supervision and then transfers the learned cross-domain representations to precise small-target detection. By grounding infrared representations in language semantics, JinSight enables a single model to generalize across heterogeneous infrared domains. We then introduce Latent Semantic Interaction (LSI), which exchanges language-aligned global semantics with fine-grained spatial features in a compact low-rank space. To address the lack of multimodal omni-domain IRST benchmarks, we build \textbf{OmniIRST-VL}, the first large-scale, highly diverse vision--language dataset for omni-domain IRST detection. It comprises over 39k annotations across six complementary instruction tasks covering both scene-level understanding and target-centric reasoning.

红外检测视觉语言跨域泛化小目标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。