arXiv:2508.07630cs.CLcs.AI2025-08中稿 · IJCNLP-AACL 2025被引 5

评测视觉语言模型跨图表推理能力,揭示其在复杂场景下的短板。

InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information

  • 构建三阶难度的跨图表推理基准,涵盖事实、整合与语义推理。
  • 模型在真实复杂图表上准确率大幅下降,尤其在多图整合时表现差。
  • 适合研究多模态推理、数据可视化分析的学者和开发者参考。

我们提出InterChart,一个诊断性基准,用于评估视觉语言模型(VLMs)在多个相关图表间进行推理的能力,该任务在科学报告、金融分析和公共政策仪表盘等实际应用中至关重要。与以往聚焦孤立、视觉统一图表的基准不同,InterChart通过多样化的题目类型——从实体推断、趋势关联到数值估算和基于2-3个主题或结构相关的图表的抽象多步推理——挑战模型。基准分为三个难度递增的层级:(1) 单个图表的事实推理;(2) 合成对齐图表集的整合分析;(3) 视觉复杂的现实世界图表对的语义推理。对当前最先进的开源与闭源VLMs的评估显示,随着图表复杂性增加,准确率持续且显著下降。我们发现,当将多实体图表分解为更简单的视觉单元时,模型表现更好,凸显其在跨图表整合方面的困难。通过暴露这些系统性局限,InterChart为复杂多视觉环境中的多模态推理发展提供了严谨框架。

原文摘要 · Abstract (English)

We introduce InterChart, a diagnostic benchmark that evaluates how well vision-language models (VLMs) reason across multiple related charts, a task central to real-world applications such as scientific reporting, financial analysis, and public policy dashboards. Unlike prior benchmarks focusing on isolated, visually uniform charts, InterChart challenges models with diverse question types ranging from entity inference and trend correlation to numerical estimation and abstract multi-step reasoning grounded in 2-3 thematically or structurally related charts. We organize the benchmark into three tiers of increasing difficulty: (1) factual reasoning over individual charts, (2) integrative analysis across synthetically aligned chart sets, and (3) semantic inference over visually complex, real-world chart pairs. Our evaluation of state-of-the-art open- and closed-source VLMs reveals consistent and steep accuracy declines as chart complexity increases. We find that models perform better when we decompose multi-entity charts into simpler visual units, underscoring their struggles with cross-chart integration. By exposing these systematic limitations, InterChart provides a rigorous framework for advancing multimodal reasoning in complex, multi-visual environments.

多模态推理图表理解视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。