用大模型先验增强特征空间,提升无源目标检测的物体聚焦能力
Foundation Model Priors Enhance Object Focus in Feature Space for Source-Free Object Detection
- 利用大模型生成的通用掩码引导网络关注物体区域
- 在严重前景背景不平衡下仍保持稳定伪标签性能
- 适合需要鲁棒性检测的工业场景应用
当前主流的无源目标检测(SFOD)方法依赖均值教师自标注,但领域偏移会削弱检测器对物体的聚焦表示,导致对背景杂乱区域产生高置信度激活,进而生成不可靠的伪标签。现有工作多聚焦于优化伪标签,却忽视了特征空间本身需要强化。本文提出FALCON-SFOD框架,包含两个互补组件:SPAR(空间先验感知正则化)利用视觉基础模型(如OV-SAM)生成的类无关二值掩码,引导网络在特征空间中形成结构化且以前景为中心的激活;IRPL(不平衡感知噪声鲁棒伪标注)则在严重前景-背景不平衡条件下实现平衡且抗噪的学习。基于理论分析,该设计可获得更紧的定位与分类误差界。FALCON-SFOD在多个SFOD基准上表现优异,验证了其有效性。
原文摘要 · Abstract (English)
Current state-of-the-art approaches in Source-Free Object Detection (SFOD) typically rely on Mean-Teacher self-labeling. However, domain shift often reduces the detector's ability to maintain strong object-focused representations, causing high-confidence activations over background clutter. This weak object focus results in unreliable pseudo-labels from the detection head. While prior works mainly refine these pseudo-labels, they overlook the underlying need to strengthen the feature space itself. We propose FALCON-SFOD (Foundation-Aligned Learning with Clutter suppression and Noise robustness), a framework designed to enhance object-focused adaptation under domain shift. It consists of two complementary components. SPAR (Spatial Prior-Aware Regularization) leverages the generalization strength of vision foundation models to regularize the detector's feature space. Using class-agnostic binary masks derived from OV-SAM, SPAR promotes structured and foreground-focused activations by guiding the network toward object regions. IRPL (Imbalance-aware Noise Robust Pseudo-Labeling) complements SPAR by promoting balanced and noise-tolerant learning under severe foreground-background imbalance. Guided by a theoretical analysis that connects these designs to tighter localization and classification error bounds, FALCON-SFOD achieves competitive performance across SFOD benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。