arXiv:2608.05771cs.CVcs.AI2026-08

提出新模型提升红外小目标跨域检测能力

HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection

论文配图:HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection
图 1 · 摘自论文原文
  • 通过扰动目标或背景扩展关系模式,增强泛化能力
  • 在多个数据集上跨域检测准确率优于现有方法
  • 适合需要强跨域适应的红外目标检测场景

红外小目标检测(IRSTD)在同域评估下已取得显著进展,但在泛化到未见红外域时性能常大幅下降。现有方法主要通过增强目标响应和抑制背景干扰来提升检测效果,但仅在有限源域上训练时,其学习到的判别规则受限于源域中的目标-背景关系模式。本文将此问题定义为目标-背景关系漂移:未见域可能呈现训练中未观察到的关系模式,从而削弱源域学得的判别能力。为此,提出HyTBE——一种双曲目标-背景专家模型,通过显式关系线索扩展源域关系模式并自适应调整视觉表征。目标-背景关系干预模块选择性扰动目标或背景,扩大训练中可观察的关系模式,同时保持有效监督;双曲关系建模模块将多尺度视觉特征映射至庞加莱球,根据特征项相对于目标与背景锚点的相对距离刻画其关系;双曲引导的MoE适配器则利用这些双曲关系表示校准多尺度特征,并对不同关系模式聚合专家特异性修正。在NUAA-SIRST、NUDT-SIRST和IRSTD-1K上的留一域外实验表明,HyTBE在跨域泛化能力上优于多种竞争基线。

原文摘要 · Abstract (English)

Infrared small target detection (IRSTD) has achieved substantial progress under domain-consistent evaluation, yet detector performance often degrades markedly when generalizing to unseen infrared domains. Existing methods primarily improve detection by enhancing target responses and suppressing background interference. However, when trained on only a limited set of source domains, their learned decision rules are inevitably established from a restricted range of source-domain target-background relation patterns. We formulate this cross-domain failure as target-background relation shift: unseen domains may exhibit relation patterns that are not observed during training, thereby weakening the discriminative capability learned from the source domains. To address this problem, we propose HyTBE, a Hyperbolic Target-Background Expert model that expands source-domain relation patterns and adaptively adjusts visual representations using explicit relation cues. The Target-Background Relation Intervention selectively perturbs either targets or backgrounds, broadening the observable relation patterns during training while maintaining valid supervision. Subsequently, the Hyperbolic Relation Modeling maps multi-scale visual cues into a Poincaré ball and characterizes the target-background relation of each feature token according to its relative distances to the target and background anchors. The Hyperbolic-guided MoE Adapter further uses these hyperbolic relation representations to calibrate multi-scale visual features and aggregate expert-specific feature corrections for different relation patterns. Leave-one-domain-out experiments on NUAA-SIRST, NUDT-SIRST, and IRSTD-1K demonstrate that HyTBE achieves stronger cross-domain generalization than competitive baselines.

红外检测跨域泛化双曲几何目标检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。