通过分层分析噪声对开放词汇目标检测器的影响,揭示了鲁棒性下降的根源。
Robust Onion: Peeling Open Vocab Object Detectors Under Noise

- 用可控合成降级逐层剖析检测器,发现特征崩溃是核心问题。
- 相同视觉主干模型在相似层出现类似特征崩溃,导致鲁棒性相近。
- 轻量级插件方案仅用96分之一参数即可提升真实数据集鲁棒性。
真实世界噪声对开放词汇目标检测器(OV-ODs)的影响因架构复杂而难以理解。我们提出全面分析 Robust Onion,通过受控的合成视觉退化,逐层剖析 OV-ODs,揭示其鲁棒性如何、为何以及在何处退化,并系统分析特征崩溃现象。研究发现,具有相似视觉主干的模型表现出相近鲁棒性,源于在相似层级发生类似的特征崩溃;而预训练策略、架构细节和描述监督等因素影响较小。鲁棒性主要由图像领域决定,而非标注信息,解释了 COCO 与 LVIS 受损程度相似的原因,也说明 ODinW-13 因包含大量孤立物体而可能产生鲁棒性被高估的现象。我们通过轻量级插件式方法 NN & TK0,在真实数据集 BDD100K、WiderFace 与 VisDRONE 上验证了上述洞察,仅需端到端训练 96 分之一的可训练参数即可提升鲁棒性。同时,我们解释了先前研究中观察到的鲁棒性现象。
原文摘要 · Abstract (English)
The impact of real-world noise on Open Vocabulary Object Detectors (OV-ODs) remains poorly understood due to their architectural complexity. We present our comprehensive analysis Robust Onion, an empirical study that uses controlled synthetic visual degradations to peel OV-ODs layer-by-layer, revealing how, why, and where robustness degrades, systematically analyzing feature collapse. Our findings reveal that models with similar vision backbones exhibit comparable robustness, driven by similar feature collapse at similar layers, while factors such as pretraining strategy, architectural nuances, and caption supervision contribute little. Robustness is primarily governed by the image domain rather than annotations, explaining the similar robustness impact on COCO and LVIS, and why datasets like ODinW-13 can give an impression of inflated robustness due to large, isolated objects. Finally, we validate our insights by improving robustness on real-world BDD100K, WiderFace, and VisDRONE via our lightweight plug-and-play NN & TK0 approach, using 96x fewer trainable parameters than end-to-end training. We also explain the prior works' robustness observations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。