arXiv:2503.12797cs.CVcs.AI2025-03被引 16

解决视觉定位中专业知识利用不足的问题

KARL: Knowledge-Aware Reasoning and Reinforcement Learning for Knowledge-Intensive Visual Grounding

  • 构建知识引导推理数据,激活模型领域知识
  • 引入动态奖励调节机制,提升跨域泛化能力
  • 适合研究多模态理解与知识增强的学者

知识密集型视觉定位(KVG)要求模型使用细粒度、领域特定的实体名称进行物体定位,而非通用指代表达。尽管多模态大语言模型(MLLM)具备丰富实体知识和强通用定位能力,但在定位专业概念时往往无法有效利用这些知识,暴露出内部知识与定位预测之间的知识-定位鸿沟。为此,我们提出一种面向KVG的知识感知训练范式。该方法首先构建知识引导的推理数据,促使模型在定位过程中激活相关领域知识;随后引入KARL框架,通过模型对不同实体知识掌握程度的估计,自适应调节奖励信号。为支持系统评估,我们构建了覆盖10个领域的KVG-Bench基准,包含1.3K条精心设计的测试用例,涉及531张图像和882个实体。大量实验表明,该方法持续优于多种基线模型,在未见类别上实现显著更强的跨域泛化能力。代码、数据与模型已开源。

原文摘要 · Abstract (English)

Knowledge-Intensive Visual Grounding (KVG) requires models to localize objects using fine-grained, domain-specific entity names rather than generic referring expressions. Although Multimodal Large Language Models (MLLMs) possess rich entity knowledge and strong generic grounding capabilities, they often fail to effectively utilize such knowledge when grounding specialized concepts, revealing a knowledge-grounding gap between internal knowledge and grounding predictions. To address this challenge, we propose a knowledge-aware training paradigm for KVG. Our approach first constructs knowledge-guided reasoning data to encourage models to activate domain-relevant entity knowledge during grounding, and then introduces KARL, a Knowledge-Aware Reinforcement Learning framework that adaptively modulates reward signals according to the model's estimated knowledge mastery of different entities. To facilitate systematic evaluation, we introduce KVG-Bench, a benchmark spanning 10 domains with 1.3K curated test cases covering 531 images and 882 entities. Extensive experiments show that our approach consistently outperforms a wide range of baseline models and achieves substantially stronger cross-domain generalization on unseen categories. The data, codes, and models are released at https://github.com/thunlp/KARL.

视觉定位知识增强强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。