arXiv:2411.10252cs.CV2024-11被引 1

让语言模型与检测模型协作,提升图像中物体定位与上下文理解能力。

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning

  • 语言模型主导推理,视觉模型提供定位与分类支持,协同优化结果。
  • 在COCO数据集上多模型检测性能显著提升,定位更准、关系更合理。
  • 适合需要精准定位与语义理解的视觉任务,如智能监控、机器人导航。

多模态大语言模型(MLLMs)在图像描述任务中表现优异,但在物体精确定位方面存在不足,而传统目标检测模型虽定位准确,却常因缺乏对物体间关系的建模而导致检测结果缺乏上下文一致性。为此,我们提出视觉-语言代理(VLA),一种融合MLLM关系推理能力与传统检测器精确定位优势的协作框架。在VLA中,语言模型作为中心语言代理,与专门负责检测和分类的视觉代理协同工作:语言代理通过分析物体间的空间与上下文关系评估并优化检测结果,分类视觉代理则提供修正反馈以提升分类准确性。该协作机制显著提升了空间推理与物体定位能力,在COCO数据集上的多模型评测中均取得显著性能提升,展现出实现高精度且具上下文一致性的目标检测新范式潜力。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) excel at descriptive tasks within images but often struggle with precise object localization, a critical element for reliable visual interpretation. In contrast, traditional object detection models provide high localization accuracy but frequently generate detections lacking contextual coherence due to limited modeling of inter-object relationships. To address this fundamental limitation, we introduce the \textbf{Visual-Linguistic Agent (VLA), a collaborative framework that combines the relational reasoning strengths of MLLMs with the precise localization capabilities of traditional object detectors. In the VLA paradigm, the MLLM serves as a central Linguistic Agent, working collaboratively with specialized Vision Agents for object detection and classification. The Linguistic Agent evaluates and refines detections by reasoning over spatial and contextual relationships among objects, while the classification Vision Agent offers corrective feedback to improve classification accuracy. This collaborative approach enables VLA to significantly enhance both spatial reasoning and object localization, addressing key challenges in multimodal understanding. Extensive evaluations on the COCO dataset demonstrate substantial performance improvements across multiple detection models, highlighting VLA's potential to set a new benchmark in accurate and contextually coherent object detection.

多模态目标检测协作推理视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。