arXiv:2507.01627cs.CL2025-07ACL被引 3

用真实分析笔记构建图表问答数据集,更贴近实际决策场景。

Chart Question Answering from Real-World Analytical Narratives

  • 从可视化笔记中提取真实多视图图表与自然语言问题
  • GPT-4.1在该数据集上准确率仅69.3%,显示现有模型仍不充分
  • 适合研究真实世界图表理解与人机分析协作的学者

我们提出一个新图表问答(CQA)数据集,源自可视化笔记本。该数据集包含真实世界、多视图图表及基于分析叙事的自然语言问题。与以往基准不同,本数据反映生态有效推理流程。对先进多模态大模型的基准测试显示显著性能差距,GPT-4.1准确率为69.3%,凸显此更真实CQA设置的挑战。

原文摘要 · Abstract (English)

We present a new dataset for chart question answering (CQA) constructed from visualization notebooks. The dataset features real-world, multi-view charts paired with natural language questions grounded in analytical narratives. Unlike prior benchmarks, our data reflects ecologically valid reasoning workflows. Benchmarking state-of-the-art multimodal large language models reveals a significant performance gap, with GPT-4.1 achieving an accuracy of 69.3%, underscoring the challenges posed by this more authentic CQA setting.

图表问答多模态真实数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。