arXiv:2511.06899cs.CLcs.AI2025-11AAAI

提出树状结构评分法,精准评估多模态模型的推理过程真确性。

RPTS: Tree-Structured Reasoning Process Scoring for Faithful Multimodal Evaluation

  • 构建推理树结构,按层级动态分配权重评估每步推理真确性。
  • 在390个实例上验证,发现闭源模型推理优势明显但仍有漏洞。
  • 首次系统分析跨模态关系对推理的影响,适合研究多模态认知的学者。

大型视觉语言模型(LVLM)在多模态推理任务中表现优异,但现有评测多采用选择题或简答形式,忽略推理过程。部分评测虽关注推理,却仅在答案错误时检查,忽略了错误推理导致正确答案的情形。此外,现有方法未考虑跨模态关系的影响。为此,本文提出基于树结构的推理过程评分(RPTS),将推理步骤组织为推理树,利用其层级信息为每步分配加权真确性分数。通过动态调整权重,RPTS不仅能评估整体推理正确性,还能定位模型推理失败节点。为验证其有效性,我们构建新基准RPTS-Eval,包含374张图像和390个推理实例,每个实例含可靠的视觉-文本线索作为推理树叶节点,并定义三类跨模态关系以研究其对推理的影响。对代表性模型(如GPT4o、Llava-Next)的评测揭示了其在多模态推理中的局限性,凸显开源与闭源模型间的差异。该基准有望推动多模态推理研究发展。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) excel in multimodal reasoning and have shown impressive performance on various multimodal benchmarks. However, most of these benchmarks evaluate models primarily through multiple-choice or short-answer formats, which do not take the reasoning process into account. Although some benchmarks assess the reasoning process, their methods are often overly simplistic and only examine reasoning when answers are incorrect. This approach overlooks scenarios where flawed reasoning leads to correct answers. In addition, these benchmarks do not consider the impact of intermodal relationships on reasoning. To address this issue, we propose the Reasoning Process Tree Score (RPTS), a tree structure-based metric to assess reasoning processes. Specifically, we organize the reasoning steps into a reasoning tree and leverage its hierarchical information to assign weighted faithfulness scores to each reasoning step. By dynamically adjusting these weights, RPTS not only evaluates the overall correctness of the reasoning, but also pinpoints where the model fails in the reasoning. To validate RPTS in real-world multimodal scenarios, we construct a new benchmark, RPTS-Eval, comprising 374 images and 390 reasoning instances. Each instance includes reliable visual-textual clues that serve as leaf nodes of the reasoning tree. Furthermore, we define three types of intermodal relationships to investigate how intermodal interactions influence the reasoning process. We evaluated representative LVLMs (e.g., GPT4o, Llava-Next), uncovering their limitations in multimodal reasoning and highlighting the differences between open-source and closed-source commercial LVLMs. We believe that this benchmark will contribute to the advancement of research in the field of multimodal reasoning.

多模态推理评测基准树结构评分跨模态关系

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。