arXiv:2506.15903cs.LG2025-06被引 3

构建超27万对图文指令数据集,推动自然语言编辑矢量图研究

VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics

  • 用CLIP配图+视觉语言模型生成指令,构建大规模图文编辑数据集
  • 现有大模型在指令编辑中准确率与有效性均不理想,任务挑战大
  • 适合研究自然语言交互、图形生成与编辑的开发者与研究人员

我们提出一个大规模指令引导的矢量图像编辑数据集,包含超过27万对SVG图像与自然语言编辑指令。该数据集支持基于文本命令修改矢量图形的模型训练与评估。数据构建过程包括通过CLIP相似性进行图像配对,以及利用视觉语言模型生成指令。初步实验表明,当前先进大语言模型在生成准确且有效的编辑结果方面表现不佳,凸显该任务的难度。为促进自然语言驱动的矢量图形生成与编辑研究,本文所创建的资源已公开可用。

原文摘要 · Abstract (English)

We introduce a large-scale dataset for instruction-guided vector image editing, consisting of over 270,000 pairs of SVG images paired with natural language edit instructions. Our dataset enables training and evaluation of models that modify vector graphics based on textual commands. We describe the data collection process, including image pairing via CLIP similarity and instruction generation with vision-language models. Initial experiments with state-of-the-art large language models reveal that current methods struggle to produce accurate and valid edits, underscoring the challenge of this task. To foster research in natural language-driven vector graphic generation and editing, we make our resources created within this work publicly available.

矢量编辑自然语言数据集生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。