通过细粒度跨模态互动提升单一领域泛化检测性能
Boosting Single-domain Generalized Object Detection via Vision-Language Knowledge Interaction
- 设计细粒度文本-视觉区域交互机制,学习跨域不变特征
- 在Cityscapes-C和DWD上分别提升8.8%和7.9%的mPC
- 适合需要应对复杂场景变化的视频监控与VR/AR应用
单领域泛化目标检测(S-DGOD)旨在仅用单一源域数据训练检测器,却能在多种未见目标域中良好泛化,适用于智能视频监控、VR/AR等存在显著域偏移的多媒体场景。尽管大型视觉语言模型取得成功,现有方法仍依赖粗粒度的预训练知识(如恶劣天气图像与对应文字描述),作为隐式正则化,难以捕捉区域与物体级别的精确特征。本文提出一种新型跨模态特征学习方法,核心为跨模态与区域感知特征交互机制,通过细粒度文本与视觉特征间的动态交互,同时学习跨模态与同模态区域不变性。此外,设计了跨域候选框精修与混合策略,对齐多域区域提议位置并增强多样性,提升未见场景下的定位能力。在S-DGOD基准数据集上,本方法在Cityscapes-C和DWD上分别取得+8.8%和+7.9%的mPC提升,达到新最优性能。
原文摘要 · Abstract (English)
Single-Domain Generalized Object Detection~(S-DGOD) aims to train an object detector on a single source domain while generalizing well to diverse unseen target domains, making it suitable for multimedia applications that involve various domain shifts, such as intelligent video surveillance and VR/AR technologies. With the success of large-scale Vision-Language Models, recent S-DGOD approaches exploit pre-trained vision-language knowledge to guide invariant feature learning across visual domains. However, the utilized knowledge remains at a coarse-grained level~(e.g., the textual description of adverse weather paired with the image) and serves as an implicit regularization for guidance, struggling to learn accurate region- and object-level features in varying domains. In this work, we propose a new cross-modal feature learning method, which can capture generalized and discriminative regional features for S-DGOD tasks. The core of our method is the mechanism of Cross-modal and Region-aware Feature Interaction, which simultaneously learns both inter-modal and intra-modal regional invariance through dynamic interactions between fine-grained textual and visual features. Moreover, we design a simple but effective strategy called Cross-domain Proposal Refining and Mixing, which aligns the position of region proposals across multiple domains and diversifies them, enhancing the localization ability of detectors in unseen scenarios. Our method achieves new state-of-the-art results on S-DGOD benchmark datasets, with improvements of +8.8\%~mPC on Cityscapes-C and +7.9\%~mPC on DWD over baselines, demonstrating its efficacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。