arXiv:2606.00782cs.CV2026-06

用连续潜空间流动生成检测查询,实现零样本开放词汇检测

FlowOVD: Learning Generative Latent Flows for Zero-shot Open-vocabulary Detection

论文配图:FlowOVD: Learning Generative Latent Flows for Zero-shot Open-vocabulary Detection
图 1 · 摘自论文原文
  • 将查询生成建模为潜空间中的连续流动过程
  • 在COCO和LVIS上分别达49.5和31.5 AP,优于GroundingDINO
  • 无需额外训练数据,适合长尾分布下的开放词汇场景

开放词汇目标检测(OVD)通过大规模视觉语言预训练取得了显著进展。现有方法通常将OVD视为判别式预测问题,解码器查询要么固定不变,要么从编码器特征初始化,限制了其多样性与灵活性。本文提出一种生成视角,将解码器查询生成建模为潜空间中的连续传输过程。我们提出FlowOVD,一个基于修正流的文本条件查询生成框架,逐步将文本无关查询转化为文本引导查询。通过在基于视觉语言模型(VLM)的检测器中引入连续潜查询动态,该方法避免了启发式离散查询构造,实现了更丰富的语义对齐。无需额外训练数据,FlowOVD在COCO上达到49.5 AP,在LVIS上达到31.5 AP,分别比GroundingDINO提升+1.2 AP(+2.5%)和+4.1 AP(+15.0%)。在更具挑战性的长尾LVIS基准上的更大提升,进一步凸显了连续查询生成在开放词汇泛化中的有效性。

原文摘要 · Abstract (English)

Open-vocabulary object detection (OVD) has achieved remarkable progress through large-scale vision-language pre-training. Existing methods, however, typically formulate OVD as a discriminative prediction problem, where decoder queries are either static or initialized from encoder features, thus limiting their diversity and flexibility. In this paper, we introduce a generative perspective by modeling decoder query generation as a continuous transport process in latent space. We propose FlowOVD, a text-conditioned query generation framework based on rectified flow that progressively transforms text-agnostic queries into text-guided queries. By introducing continuous latent query dynamics into a vision-language model (VLM) based detector, our method avoids heuristic discrete query construction and enables more expressive semantic alignment for open-vocabulary detection. Without requiring additional training data, FlowOVD achieves 49.5 AP on COCO and 31.5 AP on LVIS, outperforming GroundingDINO by +1.2 AP (+2.5 %) and +4.1 AP (+15.0 %), respectively. The larger gain on the challenging long-tailed LVIS benchmark further highlights the effectiveness of continuous query generation for open-vocabulary generalization.

开放词汇检测生成模型潜空间流动零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。