arXiv:2511.18012cs.CV2025-11

用状态和场景增强原型,提升弱监督开放词汇目标检测性能

State and Scene Enhanced Prototypes for Weakly Supervised Open-Vocabulary Object Detection

  • 引入状态感知文本描述生成更丰富的语义原型
  • 结合场景上下文优化伪框与文本对齐,减少语义错配
  • 适合研究开放词汇检测与弱监督学习的学者

开放词汇目标检测(OVOD)旨在将物体识别推广至未见类别,而弱监督开放词汇目标检测(WS-OVOD)则结合边界框级标注与图像级标签。尽管近期取得进展,仍存在两大挑战:现有语义原型虽经大语言模型增强,但静态且有限,无法捕捉由不同物体状态(如猫的姿态)引起的类内视觉差异;标准伪框生成导致视觉区域提案(含上下文)与以对象为中心的文本嵌入之间存在语义错配。为此,本文提出两种互补的原型增强策略:为捕捉类内外观与状态变化,提出状态增强语义原型(SESP),通过生成状态感知文本描述(如“一只睡觉的猫”)获得更具判别性的原型;在此基础上,进一步引入场景增强伪原型(SAPP),融合上下文语义(如“猫躺在沙发上”),并采用软对齐机制促进视觉-文本表示的一致性。结合SESP与SAPP,方法显著提升了语义原型的丰富性与视觉-文本对齐效果,在多个基准上实现显著性能提升。

原文摘要 · Abstract (English)

Open-Vocabulary Object Detection (OVOD) aims to generalize object recognition to novel categories, while Weakly Supervised OVOD (WS-OVOD) extends this by combining box-level annotations with image-level labels. Despite recent progress, two critical challenges persist in this setting. First, existing semantic prototypes, even when enriched by LLMs, are static and limited, failing to capture the rich intra-class visual variations induced by different object states (e.g., a cat's pose). Second, the standard pseudo-box generation introduces a semantic mismatch between visual region proposals (which contain context) and object-centric text embeddings. To tackle these issues, we introduce two complementary prototype enhancement strategies. To capture intra-class variations in appearance and state, we propose the State-Enhanced Semantic Prototypes (SESP), which generates state-aware textual descriptions (e.g., "a sleeping cat") to capture diverse object appearances, yielding more discriminative prototypes. Building on this, we further introduce Scene-Augmented Pseudo Prototypes (SAPP) to address the semantic mismatch. SAPP incorporates contextual semantics (e.g., "cat lying on sofa") and utilizes a soft alignment mechanism to promote contextually consistent visual-textual representations. By integrating SESP and SAPP, our method effectively enhances both the richness of semantic prototypes and the visual-textual alignment, achieving notable improvements.

目标检测弱监督开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。