arXiv:2512.05091cs.CV2025-12被引 10

让AI像人一样说出看图推理的每一步,透明化视觉思考过程。

Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark

  • 要求模型输出从目标物体到推理路径的逐级定位
  • 在真实数据上测试,现有模型正确率仅40%但路径错误
  • 适合研究可解释AI、视觉推理与模型可信度的人

多模态大模型在视觉定位和视觉问答任务中表现优异,但其推理过程仍不透明,通常只输出最终答案而无法展示中间步骤或细粒度证据(如像素、位置)。这与人类通过视觉推理链自然思考的方式形成对比。为此,我们提出视觉推理追踪(VRT)任务,要求模型不仅定位目标对象,还需显式预测构成推理路径的中间对象。我们贡献:(1) 人工标注的VRT-Bench基准,用于评估视觉推理能力;(2) 一种衡量推理轨迹质量的新指标;(3) 一个用于训练的大型数据集VRT-80k。实验表明,尽管现有模型常能给出正确最终结果,却难以准确建立中间推理路径。相比之下,在VRT-80k上训练的模型在追踪推理路径方面取得显著提升。

原文摘要 · Abstract (English)

Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved performance on tasks such as visual grounding and visual question answering. However, the reasoning processes of these models remain largely opaque; they typically output only final predictions without revealing the intermediate steps or fine-grained evidence (e.g., pixels, locations) that lead to the result. This contrasts with human intelligence, which naturally operates through a chain of visual reasoning. To address this limitation, we introduce the Visual Reasoning Tracer (VRT) task, which requires models to not only localize the target object but also explicitly predict the intermediate objects that form the reasoning path. To advance research in this area, we contribute: (1) VRT-Bench, a human-annotated benchmark for evaluating visual reasoning; (2) a new metric for assessing the quality of reasoning traces; and (3) VRT-80k, a large-scale dataset for reasoning model training. Our experiments reveal that while existing models often produce the correct final output, they struggle to ground their intermediate reasoning. In contrast, models trained on VRT-80k achieve substantial improvements in tracing the reasoning path.

视觉推理可解释性多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。