arXiv:2507.21167cs.CVcs.AI2025-07被引 8

用文字加图形标注实现精准图表修改,提升编辑准确性。

ChartM$^3$: Benchmarking Chart Editing with Multimodal Instructions

  • 结合文字与视觉标记表达修改意图,解决语言描述模糊问题。
  • 构建1000个分四难度层级的图表编辑样本,支持多维度评估。
  • 适合需要精准数据可视化编辑的研究者与工程师使用。

图表是科研与工业中广泛使用的数据可视化形式。现有方法多依赖自然语言指令进行图表编辑,但语言常因模糊而难以支持精细操作。本文提出一种新型多模态图表编辑范式,用户通过自然语言配合视觉标记明确指定修改目标。为此,我们构建了ChartM³基准数据集,包含1000个样本,涵盖四个难度层级,每条样本包含(图表、代码、多模态指令)三元组。该数据集提供视觉外观与代码正确性双重评估指标,全面评测模型表现。实验发现,当前多模态大模型(如GPT-4o)在理解与执行视觉标记方面存在明显不足。为此,我们构建了包含24,000个样本的ChartM³-Train训练集,基于此微调可显著提升模型性能,证明多模态监督对实用图表编辑系统的关键作用。数据集、代码与评估工具已开源于https://github.com/MLrollIT/ChartM3。

原文摘要 · Abstract (English)

Charts are a fundamental visualization format widely used in data analysis across research and industry. While enabling users to edit charts based on high-level intentions is of great practical value, existing methods primarily rely on natural language instructions, which are often too ambiguous to support fine-grained editing. In this work, we introduce a novel paradigm for multimodal chart editing, where user intent is expressed through a combination of natural language and visual indicators that explicitly highlight the elements to be modified. To support this paradigm, we present Chart$\text{M}^3$, a new benchmark for Multimodal chart editing with Multi-level complexity and Multi-perspective evaluation. Chart$\text{M}^3$ contains 1,000 samples spanning four levels of editing difficulty. Each sample includes triplets in the form of (chart, code, multimodal instructions). To comprehensively evaluate chart editing models, Chart$\text{M}^3$ provides metrics that assess both visual appearance and code correctness. Our benchmark reveals significant limitations in current multimodal large language models (MLLMs), including GPT-4o, particularly in their ability to interpret and act on visual indicators. To address this, we construct Chart$\text{M}^3$-Train, a large-scale training set with 24,000 multimodal chart editing samples. Fine-tuning MLLMs on this dataset leads to substantial improvements, demonstrating the importance of multimodal supervision in building practical chart editing systems. Our datasets, codes, and evaluation tools are available at https://github.com/MLrollIT/ChartM3. %https://github.com/MLrollIT/ChartM3Our datasets, codes, and evaluation tools are available at https://github.com/yaolinli/VCE.

图表生成多模态数据可视化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。