arXiv:2511.00010cs.CL2025-11被引 11

用新基准评测大模型绘图能力,提出小模型也能高效生成复杂图表。

PlotCraft: Pushing the Limits of LLMs for Complex and Interactive Data Visualization

  • 构建涵盖48种图表类型的1000个挑战性任务,评估单轮与多轮交互绘图能力。
  • 小模型PlotCraftor在困难任务上性能提升超50%,媲美主流商业方案。
  • 适合数据科学、教育和自动化可视化领域研究者使用。

近期大型语言模型在代码生成方面表现出色,但在处理大规模结构化数据的复杂可视化方面仍缺乏系统评估与开发。为此,我们提出PlotCraft,一个包含1000个挑战性可视化任务的新基准,覆盖金融、科研、社会学等多个主题。该基准围绕七类高阶可视化任务设计,涵盖48种不同图表类型,首次系统评估了单轮生成与多轮迭代优化在多样化任务复杂度下的表现。对23个主流LLM在PlotCraft上的全面评估显示,其在复杂可视化任务中存在明显短板。为弥补这一差距,我们构建了SynthVis-30K,一个基于协作智能体框架生成的大规模高质量复杂可视化代码数据集。在此基础上,我们开发了PlotCraftor——一种轻量级代码生成模型,在VisEval、PandasPlotBench及自建PlotCraft基准上表现媲美领先专有方案。尤其在高难度任务中,性能提升超过50%。相关数据集、基准与代码将开源至https://github.com/Speakn0w/PlotCraft-Benchmark。

原文摘要 · Abstract (English)

Recent Large Language Models (LLMs) have demonstrated remarkable proficiency in code generation. However, their ability to create complex visualizations for scaled and structured data remains largely unevaluated and underdeveloped. To address this gap, we introduce PlotCraft, a new benchmark featuring 1k challenging visualization tasks that cover a wide range of topics, such as finance, scientific research, and sociology. The benchmark is structured around seven high-level visualization tasks and encompasses 48 distinct chart types. Crucially, it is the first to systematically evaluate both single-turn generation and multi-turn refinement across a diverse spectrum of task complexities. Our comprehensive evaluation of 23 leading LLMs on PlotCraft reveals obvious performance deficiencies in handling sophisticated visualization tasks. To bridge this performance gap, we develope SynthVis-30K, a large-scale, high-quality dataset of complex visualization code synthesized via a collaborative agent framework. Building upon this dataset, we develope PlotCraftor, a novel code generation model that achieves strong capabilities in complex data visualization with a remarkably small size. Across VisEval, PandasPlotBench, and our proposed PlotCraft, PlotCraftor shows performance comparable to that of leading proprietary approaches. Especially, on hard task, Our model achieves over 50% performance improvement. We will release the benchmark, dataset, and code at https://github.com/Speakn0w/PlotCraft-Benchmark.

可视化大模型代码生成数据科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。