提出可编辑的符号化绘图新范式,解决专业图表生成难题
VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing

- 用mxGraph XML实现符号化绘图,支持精准生成与修改
- 构建1449张跨6大领域15子领域的结构化图表数据集
- 新评估体系验证当前模型在结构保真度和指令遵循上的不足
尽管视觉语言模型(VLMs)发展迅速,但在处理专业工作流所需的结构化、可控图表任务方面仍存在显著短板。现有方法多依赖像素级生成,在概率像素空间中操作,编辑性和保真度受限。本文提出一种新的「图即代码」范式,采用mxGraph扩展标记语言(XML)实现精确的图表生成与编辑。我们构建了VCG-Bench,一个面向视觉中心的mxGraph任务统一基准。该基准包含:(1) 涵盖6个领域、15个子领域的1,449个多样化图表数据集;(2) 明确定义生成(视觉转代码)与可编辑性(代码间转换)双重范式;(3) 设计多维度评估协议,包括mxGraph执行成功率、风格一致性评分(SCS)等指标。实验表明,当前SOTA VLMs在结构保真度和指令遵从性上表现不佳,暴露出其视觉理解与推理能力的局限。
原文摘要 · Abstract (English)
Despite the rapid advancements in Vision-Language Models (VLMs), a critical gap remains in their ability to handle structured, controllable diagrammatic tasks essential for professional workflows. Existing methods predominantly rely on pixel-based synthesis, which operates in probabilistic pixel spaces and is inherently limited in editability and fidelity. Instead, we propose a new Diagram-as-Code paradigm with symbolic logic that leverages mxGraph Extensible Markup Language (XML) for precise diagram generation and editing. We present VCG-Bench, a unified benchmark for visual-centric \texttt{mxGraph} tasks. VCG-Bench comprises: (1) a taxonomized dataset of 1,449 diverse diagrams spanning 6 domains and 15 sub-domains, (2) a paradigm definition that integrates Generation (Vision-to-Code) and Editability (Code-to-Code), (3) a Tailored Evaluation Protocol employing multi-dimensional metrics such as \texttt{mxGraph} Execution Success Rate, Style Consistency Score (SCS), etc. Experimental results highlight the challenges faced by current State-of-the-Art (SOTA) VLMs in structured fidelity and instruction compliance, reflecting their vision and reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。