arXiv:2603.29852cs.GRcs.AI2026-03中稿 · EMNLP被引 3

构建首个覆盖设计全流程的SVG生成与编辑基准,推动视觉代码模型发展

VectorGym: A Multi-Task Benchmark for SVG Code Generation, Sketching and Editing

  • 提出四类真实设计任务,含手绘转SVG、复杂多步编辑等新挑战
  • 基于人类标注的高质量数据集,要求理解设计语义与意图
  • 引入评分模型评估生成质量,适合研究视觉代码与AI设计的学者

我们提出VectorGym,一个涵盖文本与草图生成、复杂编辑及视觉理解的可扩展矢量图形(SVG)综合基准。该基准填补了与专业设计流程对齐的真实、高难度评测数据的空白。包含四项任务:新提出的草图转SVG任务(VG-Sketch);包含高阶图元的多步复杂编辑数据集(VG-Edit);文本转SVG生成(VG-Text);以及SVG描述生成(VG-Cap)。不同于以往依赖合成编辑的基准,VectorGym采用专家人工标注,需理解语义与设计意图。我们还提供基于渲染奖励的多任务强化学习基线,使用GRPO与课程学习训练Qwen3-VL 8B模型,在开源模型中达到最先进水平,优于更大规模的Qwen3-VL 235B,并接近GPT-4o表现。同时引入VLM-as-a-Judge评估指标,经人类相关性验证有效。对前沿视觉语言模型的评估揭示显著性能差距,确立VectorGym作为推动视觉代码生成研究的严格框架。

原文摘要 · Abstract (English)

We introduce VectorGym, a comprehensive benchmark suite for Scalable Vector Graphics (SVG) that spans generation from text and sketches, complex editing, and visual understanding. VectorGym addresses the lack of realistic, challenging benchmarks aligned with professional design workflows. Our benchmark comprises four tasks with expert human-authored annotations: the novel Sketch2SVG task (VG-Sketch); a new SVG editing dataset (VG-Edit) featuring complex, multi-step edits with higher-order primitives; Text2SVG generation (VG-Text); and SVG captioning (VG-Cap). Unlike prior benchmarks that rely on synthetic edits, VectorGym provides gold-standard human annotations that require semantic understanding and design intent. We also provide a multi-task reinforcement learning baseline that jointly optimizes across all four tasks using rendering-based rewards. This baseline, built on GRPO with curriculum learning, trains a Qwen3-VL 8B model that achieves state-of-the-art performance among open-source models, surpassing much larger models including Qwen3-VL 235B and matching GPT-4o. We also introduce a VLM-as-a-Judge metric for SVG generation, validated through human correlation studies. Our evaluation of frontier VLMs reveals significant performance gaps, positioning VectorGym as a rigorous framework for advancing visual code generation.

SVG生成多任务学习设计AI基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。