发现目标检测模型会学习背景关联,导致跨域性能下降。
Quantifying Context Bias in Domain Adaptation for Object Detection
- 通过遮蔽背景和特征扰动分析模型对前后景的依赖
- 揭示前后景关联导致检测性能下降,且存在因果关系
- 提出新指标量化跨域上下文偏差,适合鲁棒检测研究者
目标检测领域的域适应(DAOD)对缓解训练与部署数据分布偏移导致的性能下降至关重要。然而,影响DAOD的关键因素——由前景-背景(FG-BG)关联引发的上下文偏差——仍未被充分研究。本文针对三个核心问题展开:模型训练中是否编码了FG-BG关联?该关联与检测性能是否存在因果关系?其对DAOD有何影响?通过背景遮蔽、特征扰动及类激活图(CAM)引导的do-calculus分析,我们测量准确率下降率(即降率)。进一步提出“域关联梯度”这一新指标,定义为降率与最大均值差异(MMD)的比值。系统实验表明,基于卷积的目标检测模型确实编码了FG-BG关联。结果证明,上下文偏差不仅存在,且因果性地削弱了模型在跨域场景下的泛化能力。该结论在多种模型与数据集(包括ALDI++等先进架构)上得到验证。本研究强调需在DAOD框架中显式处理上下文偏差,为构建更鲁棒、泛化能力强的目标检测系统提供关键洞见。
原文摘要 · Abstract (English)
Domain adaptation for object detection (DAOD) has become essential to counter performance degradation caused by distribution shifts between training and deployment domains. However, a critical factor influencing DAOD - context bias resulting from learned foreground-background (FG-BG) associations - has remained underexplored. We address three key questions regarding FG BG associations in object detection: are FG-BG associations encoded during the training, is there a causal relationship between FG-BG associations and detection performance, and is there an effect of FG-BG association on DAOD. To examine how models capture FG BG associations, we analyze class-wise and feature-wise performance degradation using background masking and feature perturbation, measured via change in accuracies (defined as drop rate). To explore the causal role of FG-BG associations, we apply do-calculus on FG-BG pairs guided by class activation mapping (CAM). To quantify the causal influence of FG-BG associations across domains, we propose a novel metric - domain association gradient - defined as the ratio of drop rate to maximum mean discrepancy (MMD). Through systematic experiments involving background masking, feature-level perturbations, and CAM, we reveal that convolution-based object detection models encode FG-BG associations. Our results demonstrate that context bias not only exists but causally undermines the generalization capabilities of object detection models across domains. Furthermore, we validate these findings across multiple models and datasets, including state-of-the-art architectures such as ALDI++. This study highlights the necessity of addressing context bias explicitly in DAOD frameworks, providing insights that pave the way for developing more robust and generalizable object detection systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。