arXiv:2603.27556cs.CV2026-03

解决开放词汇目标检测在分布变化下的稳定性问题

Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method

  • 设计渐进式跨模态对齐方法,分阶段优化视觉与文本特征
  • 在多个分布外数据集上实现显著性能提升,最差场景下提升12.3%
  • 适合关注模型鲁棒性与真实场景泛化能力的研究者

开放词汇目标检测虽在识别新类别上表现优异,但其效果常依赖于领域静态假设。本文重新审视该范式,发现视觉分布偏移会破坏视觉特征与文本嵌入间的潜在对齐关系。为此,提出领域泛化的开放词汇检测(DG-OVOD)评估协议,并设计渐进式域不变跨模态对齐(PICA)方法。PICA通过基于模糊性和信号强度的多级训练课程,构建质量自适应的伪词原型,结合样本可靠性与视觉一致性进行优化,增强跨领域模态对齐稳定性。实验证明,该方法在多个分布外数据集上显著提升模型鲁棒性,最差场景下精度提升达12.3%。研究揭示了开放词汇系统泛化能力与潜在跨模态空间稳定性密切相关。

原文摘要 · Abstract (English)

Open-Vocabulary Object Detection (OVOD) has achieved remarkable success in generalizing to novel categories. However, this success often rests on the implicit assumption of domain stationarity. In this work, we revisit the OVOD paradigm and study a key vulnerability: the fragile coupling between visual manifolds and textual embeddings under distribution shifts. We first formulate Domain-Generalized Open-Vocabulary Object Detection (DG-OVOD) as an evaluation protocol for open-vocabulary recognition under visual shifts. Through empirical analysis, we observe that visual shifts can destabilize the latent cross-modal space, causing novel-category visual signals to drift away from their semantic anchors. Motivated by these observations, we propose Progressive Domain-invariant Cross-modal Alignment (PICA). PICA departs from uniform training by introducing a multi-level curriculum based on ambiguity and signal strength. It constructs a quality-adjusted curriculum over pseudo-word prototypes, refined by sample reliability and visual consistency, to encourage more stable cross-domain modality alignment. Our findings suggest that OVOD robustness under domain shifts is closely linked to the stability of the latent cross-modal alignment space. Our work provides a DG-OVOD evaluation protocol and a practical perspective on building more generalizable open-vocabulary systems beyond static laboratory conditions.

开放词汇检测跨模态对齐域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。