arXiv:2606.23354cs.CV2026-06中稿 · ICIP 2026

用可学习的符号指针提升视觉推理的可解释性,避免模型胡编坐标。

Faithful Grounded Visual Reasoning via Learned Proxy-Tokens

论文配图:Faithful Grounded Visual Reasoning via Learned Proxy-Tokens
图 1 · 摘自论文原文
  • 引入可学习的代理标记(proxy-tokens)作为图像特征的离散符号指针。
  • 在基准测试中,视觉定位准确率提升9.0个百分点,答案正确率相当。
  • 适合关注模型可解释性与可信AI的研究者和开发者。

多模态大语言模型在视觉问答任务中表现优异,但其黑箱特性限制了在关键领域的应用。基于视觉推理的方法通过将文本推理链与视觉定位信息关联来提升可解释性,通常依赖文本坐标,缺乏与视觉特征的可学习语义联系,导致语义-空间脱节,模型常生成与图像证据不符的坐标。本文提出Composer模型,采用一种基于可学习代理标记的新视觉定位机制,这些离散符号指针显式索引图像隐空间,使模型能以可操作、语义可操控的方式处理视觉区域。为严格验证该机制,我们构建了ComposerGCoT数据集,用于全面评估推理一致性与定位准确性。实验结果表明,Composer在最终答案准确率上与基于坐标的模型持平,同时视觉定位准确率提升9.0点。实证显示,离散代理标记比传统文本坐标更有效地捕捉空间语义,证明具有可学习语义链接的视觉定位机制是实现可信多模态大模型的可行路径。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have achieved remarkable success in Visual Question Answering (VQA), yet their "black-box" nature hinders deployment in critical domains. Grounded Visual Reasoning (GVR) approaches attempt to improve interpretability by explicitly couple textual rationales with visual grounding information, which are typically textual coordinates. This mechanism lacks a learnable semantic link to the visual features, often resulting in a semantic-spatial gap where the model hallucinates coordinates that do not correspond to image evidences. In this work, we introduce Composer, a MLLM that leverages a novel visual grounding mechanism based on learned proxy-tokens to promote faithful interpretability. These discrete symbolic pointers explicitly index the image latent space, allowing the model to manipulate visual regions as addressable, semantically manipulable sets. To rigorously validate our novel grounding mechanism, we constructed ComposerGCoT, a dataset synthesized to enable holistic assessment of reasoning consistency and grounding accuracy. Experimental results indicate that Composer achieves performance parity with its coordinate-based counterpart in final answer accuracy, while improving visual grounding accuracy by +9.0 points. By demonstrating that discrete proxy-tokens capture spatial semantics more effectively than typical textual coordinates, we establish that visual grounding mechanisms with learnable semantic links represent a promising path toward trustworthy and reliable MLLMs.

视觉推理可解释性代理标记多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。