构建科学图表的结构化逻辑推理评测集,可诊断模型每一步推理对错。
RealCQA-V2: A Diagnostic Benchmark for Structured Visual Entailment over Scientific Charts
- 将图表问答转为视觉前提证明任务,拆解每步推理到具体图表元素
- 提出链级指标,可评估完整逻辑正确性及部分推理进展
- 发现主流模型存在局部正确但整体不连贯的推理缺陷
多模态推理模型常生成流畅且看似合理的解释,但现有评测仅关注最终答案正确性,无法验证中间步骤的视觉蕴含关系,尤其缺乏对视觉组合逻辑的细粒度检验。在科学图表理解中,答案依赖于坐标轴、图例、数据标记等具有确定语义的视觉元素。我们提出 RealCQA-V2,一个大规模基准,将图表问答重构为视觉前提证明(VPP):基于图表元素的结构化逻辑蕴含任务。每个问题被拆解为人工标注的原子前提,源自坐标轴、图例、数据点及数量关系,形成可执行的推理链,而非自由文本解释。这些前提构成组合式推理链,支持对单个视觉陈述及完整推理序列的验证。我们引入链级指标,衡量逻辑完整性(AccVPP)和失败链中的推理进度(DCP),超越传统VQA准确率。对代表性大视觉语言模型的基线评估显示一致的局部-全局推理差距:模型虽能正确验证多数前提,却难以保持整条链的逻辑连贯性。RealCQA-V2为真实科学图表上的结构化视觉蕴含提供了可复现的评测基准,推动多模态推理诊断从仅看答案迈向全过程分析。
原文摘要 · Abstract (English)
Multimodal reasoning models often produce fluent answers supported by seemingly coherent rationales. Existing benchmarks evaluate only final-answer correctness. They do not support atomic visual entailment verification of intermediate steps, especially visual compositional logic. This limitation is especially acute in scientific chart understanding, where answers depend on deterministically grounded visual semantics such as axes, legends, and quantitative relations. We introduce RealCQA-V2, a large-scale benchmark that reformulates chart question answering as Visual Premise Proving (VPP): a structured logical entailment task over chart-grounded visual predicates. Each question is deconstructed into manually curated, atomic premises grounded in chart elements (axes, legends, marks, and quantitative relations), yielding executable reasoning chains rather than free-form textual rationales. These premises form compositional reasoning chains, enabling verification at the level of individual visual statements and complete reasoning sequences. We introduce chain-level metrics that measure both full logical validity (AccVPP) and partial reasoning progress within failed chains (DCP), extending beyond traditional VQA accuracy. Baseline evaluations across representative LVLMs reveal a consistent local-global reasoning gap: models often verify many individual premises correctly while failing to preserve coherence across the full chain. RealCQA-V2 establishes a reproducible benchmark for structured visual entailment over real scientific charts and enables rigorous diagnosis of multimodal reasoning beyond answer-only evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。