用因果视觉提示提升单一源域泛化目标检测能力
Towards Single-Source Domain Generalized Object Detection via Causal Visual Prompts
- 引入跨注意力提示模块,减少对颜色等虚假特征的依赖
- 双分支适配器分离因果与虚假特征,提升域适应效果
- 在复杂干扰下表现更稳健,适合真实场景部署
单源域泛化目标检测(SDGOD)是计算机视觉前沿课题,旨在通过单一源域训练提升模型在未见目标域的泛化能力。现有主流方法依赖数据增强缓解域差异,但受限于域偏移和领域知识不足,模型易陷入虚假相关性陷阱,过度依赖颜色等简单分类特征,而非物体轮廓等域不变表示。为此,我们提出Cauvis(因果视觉提示)方法:首先设计跨注意力提示模块,通过视觉提示与跨注意力机制抑制虚假特征偏差;其次提出双分支适配器,通过高频特征提取实现因果-虚假特征解耦,同时完成域适应。Cauvis在SDGOD数据集上相较现有方法取得15.9%-31.4%的性能提升,且在复杂干扰环境下表现出显著鲁棒性优势。
原文摘要 · Abstract (English)
Single-source Domain Generalized Object Detection (SDGOD), as a cutting-edge research topic in computer vision, aims to enhance model generalization capability in unseen target domains through single-source domain training. Current mainstream approaches attempt to mitigate domain discrepancies via data augmentation techniques. However, due to domain shift and limited domain-specific knowledge, models tend to fall into the pitfall of spurious correlations. This manifests as the model's over-reliance on simplistic classification features (e.g., color) rather than essential domain-invariant representations like object contours. To address this critical challenge, we propose the Cauvis (Causal Visual Prompts) method. First, we introduce a Cross-Attention Prompts module that mitigates bias from spurious features by integrating visual prompts with cross-attention. To address the inadequate domain knowledge coverage and spurious feature entanglement in visual prompts for single-domain generalization, we propose a dual-branch adapter that disentangles causal-spurious features while achieving domain adaptation via high-frequency feature extraction. Cauvis achieves state-of-the-art performance with 15.9-31.4% gains over existing domain generalization methods on SDGOD datasets, while exhibiting significant robustness advantages in complex interference environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。