新基准KnowDR-REC测试模型用真实知识推理图像目标的能力
KnowDR-REC: A Benchmark for Referring Expression Comprehension with Real-World Knowledge
- 基于真实世界知识构建,需跨模态细粒度推理
- 引入精细编辑的负样本,评估模型抗幻觉能力
- 提出新评测指标,揭示模型推理机制与缺陷
指代表达理解(REC)是多模态任务,旨在根据文本描述准确识别图像中的目标对象。然而,传统基准依赖图像内线索或缺乏细粒度实例标注,难以评估多模态大模型(MLLMs)的推理能力。为此,我们提出新基准KnowDR-REC,具备三大特征:首先,基于真实世界知识,要求细粒度跨模态推理;其次,通过精细文本编辑构建负样本,评估模型鲁棒性与抗幻觉能力;最后,引入三项新评测指标,系统分析模型内部推理过程。我们在KnowDR-REC上评估16个先进多模态模型,结果显示现有MLLMs在知识驱动的视觉定位任务中仍表现不足。进一步发现,文本理解与视觉定位在MLLMs中存在解耦现象,许多模型受记忆化捷径关联影响,严重干扰其在本基准上的表现,阻碍真正多模态推理。我们期望该基准能推动未来研究发展更鲁棒、可解释、知识密集型的视觉定位框架,提升复杂现实场景下多模态系统的可靠性。
原文摘要 · Abstract (English)
Referring Expression Comprehension (REC) is a popular multimodal task that aims to accurately detect target objects within a single image based on a given textual expression. However, due to the limitations of earlier models, traditional REC benchmarks either rely solely on intra-image cues or lack sufficiently fine-grained instance annotations, making them inadequate for evaluating the reasoning capabilities of Multi-modal Large Language Models (MLLMs). To address this gap, we propose a new benchmark, KnowDR-REC, characterized by three key features: Firstly, it is built upon real-world knowledge, requiring fine-grained multimodal reasoning across text and image. Secondly, the dataset includes elaborately constructed negative samples via fine-grained expression editing, designed to evaluate a model's robustness and anti-hallucination ability. Lastly, we introduce three novel evaluation metrics to systematically explore the model's internal reasoning process. We evaluate 16 state-of-the-art multimodal models on KnowDR-REC, with experimental results showing that existing MLLMs still struggle with knowledge-driven visual grounding tasks. Furthermore, we observe a decoupling between textual understanding and visual grounding in MLLMs, where many models are significantly influenced by memorized shortcut correlations, which severely affect their behavior on our benchmark and hinder genuine multimodal reasoning. We anticipate that the proposed benchmark will inspire future research towards developing more robust, interpretable, and knowledge-intensive visual grounding frameworks, driving the development of more reliable and robust multimodal systems for complex real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。