arXiv:2509.04457cs.CL2025-09被引 4

测试大模型真能看懂图表吗?发现它们常靠猜,不靠推理。

Do MLLMs Really Understand the Charts?

  • 设计新评测集ChartVRBench,专测图表推理能力
  • 新方法让模型在无标注图上准确率提升37.2%
  • 适合想提升图表理解能力的研究者和工程师

尽管多模态大语言模型在图表理解任务中表现日益出色,但其在处理无标注图表时仍存在严重幻觉与性能下降。本文认为现有模型主要依赖视觉识别而非视觉推理,而数值估算正是图表理解中最基础的复杂视觉推理能力。为此,我们提出ChartVRBench评测基准,专门用于隔离并评估模型的视觉推理能力。同时,我们提出基于新型视觉推理强化微调(VR-RFT)策略训练的ChartVR-3B/7B模型,以增强真实图表推理能力。大量实验表明,ChartVR在ChartVRBench上表现优于多个主流闭源模型,且其培养的视觉推理能力具有强泛化性,在多个公开图表理解基准上均取得显著提升。代码与数据集将在发表后公开。

原文摘要 · Abstract (English)

Although Multimodal Large Language Models (MLLMs) have demonstrated increasingly impressive performance in chart understanding, most of them exhibit alarming hallucinations and significant performance degradation when handling non-annotated charts. We argue that current MLLMs rely largely on visual recognition rather than visual reasoning to interpret the charts, and visual estimation of numerical values is one of the most fundamental capabilities in chart understanding that require complex visual reasoning. To prove this, we introduce ChartVRBench, a benchmark meticulously designed to isolate and evaluate visual reasoning ability in chart understanding. Furthermore, we propose ChartVR-3B/7B trained with a novel Visual Reasoning Reinforcement Finetuning (VR-RFT) strategy to strengthen genuine chart visual reasoning abilities. Extensive experiments show that ChartVR achieves superior performance on ChartVRBench, outperforming even powerful proprietary models. Moreover, the visual reasoning skills cultivated by the proposed VR-RFT demonstrate strong generalization, leading to significant performance gains across a diverse suite of public chart understanding benchmarks. The code and dataset will be publicly available upon publication.

图表理解视觉推理多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。