arXiv:2505.23242cs.CL2025-05EMNLP被引 6

构建真实复杂图表问答新基准,提升模型理解力。

ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answering

  • 提出兼顾多语言与开放输出的综合评测框架。
  • 在7类任务上验证,显著优于传统三类方法。
  • 适合研究多模态推理与实际图表分析的学者。

图表问答(CQA)已成为评估视觉-语言模型推理能力的关键多模态任务。尽管早期方法通过关注视觉特征或利用大规模预训练取得了良好表现,但现有评估多依赖固定输出格式和客观指标,忽视了真实场景下复杂的图表分析需求。本文提出ChartMind,一个面向真实世界复杂CQA任务的新基准。该基准涵盖7个任务类别,支持多语言上下文、开放域文本输出及多样图表格式,弥合了实际应用与传统学术评测之间的差距。此外,我们设计了一种上下文感知且模型无关的框架ChartLLM,聚焦提取关键上下文要素,降低噪声,提升多模态大模型的推理准确性。在ChartMind及三个代表性公开基准上,对14个主流多模态模型的广泛评估表明,该框架显著优于三种常见CQA范式:指令遵循、OCR增强与思维链,凸显灵活图表理解在真实场景下的重要性。研究为未来更鲁棒的图表推理提供了新方向。

原文摘要 · Abstract (English)

Chart question answering (CQA) has become a critical multimodal task for evaluating the reasoning capabilities of vision-language models. While early approaches have shown promising performance by focusing on visual features or leveraging large-scale pre-training, most existing evaluations rely on rigid output formats and objective metrics, thus ignoring the complex, real-world demands of practical chart analysis. In this paper, we introduce ChartMind, a new benchmark designed for complex CQA tasks in real-world settings. ChartMind covers seven task categories, incorporates multilingual contexts, supports open-domain textual outputs, and accommodates diverse chart formats, bridging the gap between real-world applications and traditional academic benchmarks. Furthermore, we propose a context-aware yet model-agnostic framework, ChartLLM, that focuses on extracting key contextual elements, reducing noise, and enhancing the reasoning accuracy of multimodal large language models. Extensive evaluations on ChartMind and three representative public benchmarks with 14 mainstream multimodal models show our framework significantly outperforms the previous three common CQA paradigms: instruction-following, OCR-enhanced, and chain-of-thought, highlighting the importance of flexible chart understanding for real-world CQA. These findings suggest new directions for developing more robust chart reasoning in future research.

图表问答多模态基准测试大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。