arXiv:2608.01328cs.CLcs.AI2026-08

构建复杂多图表推理评测基准,揭示大模型在多图任务中的性能瓶颈

LongChart VQA: A Comprehensive Benchmark for MLLMs with Complex Multi-Chart Reasoning

论文配图:LongChart VQA: A Comprehensive Benchmark for MLLMs with Complex Multi-Chart Reasoning
图 1 · 摘自论文原文
  • 基于潜在图的合成流程生成包含6.5张图、31.2个问题的多图表数据集
  • 10个顶尖模型在复杂推理中准确率显著下降,且波动大
  • 适合研究多图表理解、模型鲁棒性及推理机制的学者使用

多模态大语言模型(MLLMs)正快速发展,具备更长上下文和更强推理能力,能够处理多图表理解与多步推理。这类能力在复杂代理任务中日益重要。然而,现有基准大多侧重单图表感知,简单图表间连接不足以评估其能力。为捕捉多图表复杂性并确保一致性和有效性,我们设计了基于潜在图的合成流程。在此基础上,提出LongChart基准,其视觉问答集平均包含6.5张图像和31.2个问题。我们评估了10个最先进MLLMs,分析了推理模式、辅助工具和对图像扰动的鲁棒性三个因素的影响。结果表明,随着计算复杂度增加,模型准确率显著下降且波动剧烈,揭示了未来多图表推理研究的重要方向。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) are rapidly evolving with expanded context windows and stronger reasoning capabilities, enabling multi-chart understanding and multi-step inference. These abilities are increasingly important as MLLMs are adopted in complex agentic tasks. However, existing benchmarks largely emphasize single-chart perception, while simple chart-to-chart connections are insufficient to evaluate these capabilities. To capture multi-chart complexity while ensuring consistency and validity, we design a synthesis pipeline supported by latent graphs. Building on this pipeline, we introduce LongChart, a benchmark whose VQA sets contain an average of 6.5 images and 31.2 questions. We evaluate 10 state-of-the-art MLLMs and examine three factors that influence performance: reasoning patterns, auxiliary tools, and robustness to image perturbations. Our results show that MLLM accuracy decreases and varies substantially as computational complexity increases, highlighting directions for future research in multi-chart reasoning.

多图表推理评测基准大模型评估视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。