arXiv:2511.21064cs.AIcs.CV2025-11

让目标检测主动推理并自我进化,提升罕见物体识别能力。

OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection

  • 用马尔可夫-赌博机制实现视觉推理与自适应检测
  • 在COCO和LVIS上显著提升稀有类别检测性能
  • 适合研究开放词汇检测与智能推理系统的人看

开放词汇目标检测(OVOD)旨在通过语义信息实现跨类别泛化。尽管现有方法在大规模视觉语言数据集上预训练,其推理仍局限于固定类别名称,造成多模态训练与单模态推理之间的鸿沟。已有研究表明,优化文本表征可显著提升OVOD性能,表明文本空间仍有待探索。为此,我们提出OVOD-Agent,将被动类别匹配转化为主动视觉推理与自演化检测。受思维链(CoT)启发,OVOD-Agent将文本优化过程扩展为可解释的视觉思维链(Visual-CoT),包含显式动作。由于OVOD轻量特性,不适用基于大语言模型的管理;我们将其视觉上下文转换建模为八状态空间上的弱马尔可夫决策过程(w-MDP),自然表达代理的状态、记忆与交互动态。一个赌博模块在有限监督下生成探索信号,帮助代理聚焦不确定区域并调整检测策略。我们进一步将马尔可夫转移矩阵与赌博轨迹结合,实现自监督奖励模型(RM)优化,形成从赌博探索到RM学习的闭环。在COCO和LVIS上的实验表明,OVOD-Agent在各类OVOD主干网络中均带来一致改进,尤其在稀有类别上表现突出,验证了该框架的有效性。

原文摘要 · Abstract (English)

Open-Vocabulary Object Detection (OVOD) aims to enable detectors to generalize across categories by leveraging semantic information. Although existing methods are pretrained on large vision-language datasets, their inference is still limited to fixed category names, creating a gap between multimodal training and unimodal inference. Previous work has shown that improving textual representation can significantly enhance OVOD performance, indicating that the textual space is still underexplored. To this end, we propose OVOD-Agent, which transforms passive category matching into proactive visual reasoning and self-evolving detection. Inspired by the Chain-of-Thought (CoT) paradigm, OVOD-Agent extends the textual optimization process into an interpretable Visual-CoT with explicit actions. OVOD's lightweight nature makes LLM-based management unsuitable; instead, we model visual context transitions as a Weakly Markovian Decision Process (w-MDP) over eight state spaces, which naturally represents the agent's state, memory, and interaction dynamics. A Bandit module generates exploration signals under limited supervision, helping the agent focus on uncertain regions and adapt its detection policy. We further integrate Markov transition matrices with Bandit trajectories for self-supervised Reward Model (RM) optimization, forming a closed loop from Bandit exploration to RM learning. Experiments on COCO and LVIS show that OVOD-Agent provides consistent improvements across OVOD backbones, particularly on rare categories, confirming the effectiveness of the proposed framework.

目标检测视觉推理自进化开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。