arXiv:2506.13326cs.CVcs.HC2025-06被引 7

用AI自动评判和改进大模型生成的图表,效果媲美更大型模型。

VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation

  • 构建高质量图表评价数据集,融合人工与生成样本
  • 70亿参数小模型借助该数据集性能大幅提升
  • 适合想提升数据可视化生成质量的研究者与开发者

利用大语言模型(LLM)生成数据可视化已展现良好前景,但常产生需人工修正的次优结果。本文提出VIS-Shepherd,一种基于多模态大语言模型(MLLM)的专用评判系统,用于评估并反馈LLM生成的可视化内容。核心是构建高质量的可视化评价数据集:收集人工创作的可视化实例,合成对应的LLM生成版本,并构建高质量评价文本。通过模型自动评估与人类偏好实验验证方法有效性。实验表明,仅使用70亿参数的小型开源MLLM模型,借助该数据集即可实现显著性能提升,达到远大于自身规模的开源或专有模型水平。本工作展示了基于MLLM的自动化可视化评判的巨大潜力,并为改进LLM驱动的数据可视化生成指明了新方向。项目页面:https://github.com/bopan3/VIS-Shepherd。

原文摘要 · Abstract (English)

Data visualization generation using Large Language Models (LLMs) has shown promising results but often produces suboptimal visualizations that require human intervention for improvement. In this work, we introduce VIS-Shepherd, a specialized Multimodal Large Language Model (MLLM)-based critic to evaluate and provide feedback for LLM-generated data visualizations. At the core of our approach is a framework to construct a high-quality visualization critique dataset, where we collect human-created visualization instances, synthesize corresponding LLM-generated instances, and construct high-quality critiques. We conduct both model-based automatic evaluation and human preference studies to evaluate the effectiveness of our approach. Our experiments show that even small (7B parameters) open-source MLLM models achieve substantial performance gains by leveraging our high-quality visualization critique dataset, reaching levels comparable to much larger open-source or even proprietary models. Our work demonstrates significant potential for MLLM-based automated visualization critique and indicates promising directions for enhancing LLM-based data visualization generation. Our project page: https://github.com/bopan3/VIS-Shepherd.

可视化生成多模态模型自动评价大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。