arXiv:2603.25804cs.CL2026-03被引 4

用真实数据评测图表生成代码,发现大模型表现大幅下降。

RealChart2Code: Advancing Chart-to-Code Generation with Real Data and Multi-Task Evaluation

  • 基于2800个真实数据图表构建新基准,支持多轮对话迭代优化。
  • 14个主流视觉语言模型在复杂多面板图表上平均准确率不足50%。
  • 首次揭示开源与闭源模型在真实场景下的显著性能差距。

视觉语言模型(VLMs)在多个领域展现出强大的代码生成能力,但其对真实世界数据中复杂多面板可视化图表的复现能力仍缺乏系统评估。为此,我们提出 exttt{RealChart2Code},一个包含超过2800个实例的大规模基准,数据源自真实数据集,任务具有明确分析意图。该基准是首个系统评估从大规模原始数据生成图表并支持多轮对话式代码迭代优化的评测体系。对14个领先VLMs在 exttt{RealChart2Code} 上的全面评估显示,相比简单基准,其性能显著下降,暴露出在复杂图表结构和真实数据处理上的严重不足。分析发现,闭源模型与开源模型间存在显著性能差距,即使最先进模型也常无法准确还原复杂的多面板图表。这些结果揭示了当前VLMs的核心局限,为未来研究指明方向。相关数据与代码已开源于 https://github.com/Speakn0w/RealChart2Code。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have demonstrated impressive capabilities in code generation across various domains. However, their ability to replicate complex, multi-panel visualizations from real-world data remains largely unassessed. To address this gap, we introduce \textbf{\texttt{RealChart2Code}}, a new large-scale benchmark with over 2,800 instances grounded in authentic datasets and featuring tasks with clear analytical intent. Crucially, it is the first benchmark to systematically evaluate chart generation from large-scale raw data and assess iterative code refinement in a multi-turn conversational setting. Our comprehensive evaluation of 14 leading VLMs on \texttt{RealChart2Code} reveals significant performance degradation compared to simpler benchmarks, highlighting their struggles with complex plot structures and authentic data. Our analysis uncovers a substantial performance gap between proprietary and open-weight models and confirms that even state-of-the-art VLMs often fail to accurately replicate intricate, multi-panel charts. These findings provide valuable insights into the current limitations of VLMs and guide future research directions. We release the benchmark and code at \url{https://github.com/Speakn0w/RealChart2Code}.

图表生成视觉语言模型多轮对话真实数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。