arXiv:2604.21502cs.CV2026-04

用视觉大模型提升检测器在极端环境下的稳定性。

VFM$^{4}$SDG: Unveiling the Power of VFMs for Single-Domain Generalized Object Detection

论文配图:VFM$^{4}$SDG: Unveiling the Power of VFMs for Single-Domain Generalized Object Detection
图 1 · 摘自论文原文
  • 引入冻结的视觉大模型作为跨域关系先验,增强编码器与解码器的稳定性。
  • 在多个基准上显著优于现有方法,尤其在恶劣环境下性能提升明显。
  • 适合关注鲁棒目标检测、尤其是基于DETR架构的研究者使用。

现实世界中的天气、光照和成像变化常导致严重域偏移,使单源检测器在未见环境中性能下降。现有单域广义目标检测(SDGOD)方法主要依赖数据增强或域不变学习,却忽视了域偏移对检测器预测稳定性的影响。分析实验表明,性能下降主要源于漏检增多。进一步分析发现,这是由于域偏移破坏了DETR类检测器中编码器侧的对象-背景及实例间关系的跨域稳定性,进而削弱了解码器查询与真实对象之间的语义-空间绑定。受此启发,我们发现视觉基础模型(VFMs)在严重域偏移下仍能保持稳定的结构关系和对象响应,适合作为跨域稳定性先验来补偿检测器退化。为此,我们提出VFM⁴SDG,一种双先验学习框架,将冻结的VFM引入编码器表征学习和解码器查询建模。具体地,提出跨域稳定关系先验蒸馏,从VFM中蒸馏出稳定对象-背景与实例间关系到编码器,以缓解关系退化;同时提出基于语义-上下文先验的查询增强,将类别语义原型与全局对象上下文注入解码器前的查询,增强语义-空间绑定稳定性。大量实验表明,VFM⁴SDG在标准SDGOD基准及两种主流DETR检测框架上均显著优于现有先进方法,验证了其有效性、鲁棒性与通用性。

原文摘要 · Abstract (English)

Real-world weather, illumination, and imaging variations often induce severe domain shifts, degrading single-source detectors in unseen environments. Existing single-domain generalized object detection (SDGOD) methods mainly rely on data augmentation or domain-invariant learning, while largely overlooking how domain shift disrupts detector prediction stability. Through analytical experiments, we find that performance degradation is mainly dominated by increasing missed detections. Further analysis shows that this phenomenon stems from reduced cross-domain stability in DETR-style detectors: domain shift disrupts encoder-side object-background and inter-instance relations, and further weakens the semantic-spatial binding between decoder queries and real objects. Motivated by this, we find that vision foundation models (VFMs) still preserve stable relational structures and object responses under severe shifts, making them suitable cross-domain stability priors to compensate for detector degradation. To this end, we propose VFM$^{4}$SDG, a dual-prior learning framework for SDGOD, which introduces a frozen VFM into encoder representation learning and decoder query modeling. Specifically, we propose Cross-domain Stable Relational Prior Distillation to distill stable object-background and inter-instance relations from the VFM into the encoder, compensating for relational degradation. Meanwhile, we propose Semantic-Contextual Prior-based Query Enhancement, which injects category semantic prototypes and global object context into queries before they enter the decoder layer, enhancing semantic-spatial query-object binding stability. Extensive experiments show that VFM$^{4}$SDG significantly outperforms existing advanced methods on standard SDGOD benchmarks and two mainstream DETR-based detection frameworks, demonstrating its effectiveness, robustness, and generality.

目标检测域泛化视觉大模型DETR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。