arXiv:2503.04983cs.CV2025-03综述被引 6

用大模型提升矢量图生成与编辑效率,解决传统方法慢且复杂的问题。

Leveraging Large Language Models For Scalable Vector Graphics Processing: A Review

  • 利用大模型处理文本化的矢量图(SVG),实现高效生成与编辑。
  • 增强型模型在生成和理解任务中显著优于普通大模型。
  • 适合关注矢量设计自动化、AI辅助创意工作的研究人员与开发者。

近年来,计算机视觉的快速发展显著提升了位图图像的处理与生成能力,但作为数字设计核心的矢量图形因可缩放性和易编辑性仍相对研究不足。传统矢量化技术常存在处理时间长、输出复杂度高等问题,限制了实际应用。大语言模型(LLMs)的出现为矢量图形的生成、编辑与分析带来了新可能,尤其在本质为文本的SVG格式上表现突出。本文系统综述了现有基于LLM的SVG处理方法,将其分为生成、编辑与理解三类任务。重点分析了IconShop、StrokeNUWA和StarVector等代表性模型,揭示其优势与局限。同时,评估了SVGEditBench、VGBench和SGP-Bench等基准数据集,并通过实验发现,经过向量图形推理增强的模型在生成与理解任务中显著优于标准LLM。研究强调需构建更多样、更丰富标注的数据集以进一步提升模型性能。

原文摘要 · Abstract (English)

In recent years, rapid advances in computer vision have significantly improved the processing and generation of raster images. However, vector graphics, which is essential in digital design, due to its scalability and ease of editing, have been relatively understudied. Traditional vectorization techniques, which are often used in vector generation, suffer from long processing times and excessive output complexity, limiting their usability in practical applications. The advent of large language models (LLMs) has opened new possibilities for the generation, editing, and analysis of vector graphics, particularly in the SVG format, which is inherently text-based and well-suited for integration with LLMs. This paper provides a systematic review of existing LLM-based approaches for SVG processing, categorizing them into three main tasks: generation, editing, and understanding. We observe notable models such as IconShop, StrokeNUWA, and StarVector, highlighting their strengths and limitations. Furthermore, we analyze benchmark datasets designed for assessing SVG-related tasks, including SVGEditBench, VGBench, and SGP-Bench, and conduct a series of experiments to evaluate various LLMs in these domains. Our results demonstrate that for vector graphics reasoning-enhanced models outperform standard LLMs, particularly in generation and understanding tasks. Furthermore, our findings underscore the need to develop more diverse and richly annotated datasets to further improve LLM capabilities in vector graphics tasks.

矢量图大模型SVG生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。