arXiv:2510.11341cs.CV2025-10被引 14

用大模型统一处理SVG图文理解、编辑与生成,效果显著优于现有方法。

InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models

  • 基于多模态大模型构建统一框架,支持从静态图标到动态动画的全任务覆盖。
  • 自建SAgoge数据集含超10万样本,涵盖图标、科学图示与动态动画等多类场景。
  • 提出两阶段训练策略与专属编码方式,提升复杂结构理解与跨任务泛化能力。

通用SVG建模仍面临数据集分散、方法迁移性差及结构复杂性难以处理等问题。为此,我们利用多模态大语言模型(MLLM)强大的迁移与泛化能力,实现对SVG理解、编辑与生成的统一建模。提出InternSVG系列,包含集成的数据-基准-模型体系。核心为SAgoge,目前最大最全面的多模态SVG数据集,涵盖静态图形与动态动画,包括图标、长序列插图、科学图示与动态内容,支持多种难度任务,具有更深层级与丰富属性。基于此,引入SArena基准,提供完整任务定义与标准化评估,覆盖SAgoge所涉领域与难度范围。在此基础上,提出InternSVG,一种统一的MLLM,通过专有SVG符号、子词嵌入初始化及两阶段训练策略(从短静态SVG逐步过渡至长序列插图与复杂动画),实现正向迁移并提升整体性能。在SArena及先前基准上的实验表明,InternSVG显著优于领先开源与闭源模型。

原文摘要 · Abstract (English)

General SVG modeling remains challenging due to fragmented datasets, limited transferability of methods across tasks, and the difficulty of handling structural complexity. In response, we leverage the strong transfer and generalization capabilities of multimodal large language models (MLLMs) to achieve unified modeling for SVG understanding, editing, and generation. We present the InternSVG family, an integrated data-benchmark-model suite. At its core is SAgoge, the largest and most comprehensive multimodal dataset for SVG tasks, encompassing both static graphics and dynamic animations. It covers icons, long-sequence illustrations, scientific diagrams, and dynamic animations, supporting tasks of varied difficulty levels and providing deeper hierarchies with richer attributes compared to previous datasets. Based on this resource, we introduce SArena, a companion benchmark with comprehensive task definitions and standardized evaluation that aligns with the domains and difficulty spectrum covered by SAgoge. Building on these foundations, we propose InternSVG, a unified MLLM for SVG understanding, editing, and generation with SVG-specific special tokens, subword-based embedding initialization, and a two-stage training strategy that progresses from short static SVGs to long-sequence illustrations and complex animations. This unified formulation induces positive transfer and improves overall performance. Experiments on SArena and prior benchmark confirm that InternSVG achieves substantial gains and consistently outperforms leading open and proprietary counterparts.

SVG建模多模态大模型统一框架生成与编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。