通过反事实推理提升开放词汇目标检测在分布偏移下的鲁棒性
FACTOR: Counterfactual Training-Free Test-Time Adaptation for Open-Vocabulary Object Detection

- 基于反事实视觉属性扰动,判断预测是否受无关特征干扰
- 无需参数更新,在PASCAL-C等3个数据集上显著优于现有方法
- 适合关注模型鲁棒性与测试时自适应的开发者
开放词汇目标检测在分布偏移下常因非因果视觉属性(如亮度、纹理)与类别间的虚假关联而失效。现有测试时自适应(TTA)方法要么依赖昂贵的在线优化,要么进行全局校准,忽略了失败的属性特异性。为此,我们提出FACTOR(counterFACtual training-free Test-time adaptation for Open-vocabulaRy object detection),一个基于反事实推理的轻量级框架。通过沿非因果属性扰动测试图像,并对比原始与反事实视图下的区域级预测,FACTOR量化属性敏感度、语义相关性和预测变化,从而选择性抑制依赖属性的预测,且无需参数更新。在PASCAL-C、COCO-C和FoggyCityscapes上的实验表明,FACTOR持续优于先前的TTA方法,证明显式反事实推理能有效提升分布偏移下的鲁棒性。
原文摘要 · Abstract (English)
Open-vocabulary object detection often fails under distribution shifts, as it can be misled by spurious correlations between non-causal visual attributes (e.g., brightness, texture) and object categories. Existing test-time adaptation (TTA) methods either depend on costly online optimization or perform global calibration, overlooking the attribute-specific nature of these failures. To address this, we propose FACTOR (counterFACtual training-free Test-time adaptation for Open-vocabulaRy object detection), a lightweight framework grounded in counterfactual reasoning. By perturbing test images along non-causal attributes and comparing region-level predictions between original and counterfactual views, FACTOR quantifies attribute sensitivity, semantic relevance, and prediction variation to selectively suppress attribute-dependent predictions-without parameter updates. Experiments on PASCAL-C, COCO-C, and FoggyCityscapes show that FACTOR consistently outperforms prior TTA methods, demonstrating that explicit counterfactual reasoning effectively improves robustness under distribution shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。