为文本生成矢量图设计了基于视觉的评测框架,更贴近人眼感知。
SVGEval: A Vision-Grounded Framework for Perceptual-Quality Benchmarking and Evaluation in Text-to-SVG Generation

- 通过渲染图像评估模型输出,而非仅看SVG代码。
- 发现模型在几何和布局判断上明显落后于语义与美学评价。
- 提供可解释的评分器,适合研究与优化文本到SVG生成模型。
多模态大模型越来越多地用于生成可缩放矢量图形(SVG),但可靠的评估方法仍不充分。现有评测通常以代码为中心,或在将SVG渲染为位图后借用图像指标,无法反映人类感知,也忽略了几何和空间构图等SVG特有质量。我们提出SVGEval,一个基于视觉的多模态基准,用于人类对齐的SVG质量评估。该框架显式引入视觉渲染结果,评估模型是否能判断渲染后的实际效果,而非仅分析SVG代码,并通过多轮人工标注与专家精修获得高质量标注。对代表性多模态模型的系统评估显示:模型在语义对齐和美学方面表现较好,但在几何和布局判断上存在明显短板。基于SVGEval,我们训练了一个可解释的SVG质量评分器,输出多维度分数并附带文本推理。消融实验表明,显式的视觉锚定和推理监督对空间与几何评估尤为关键。SVGEval为多模态时代下的SVG生成评估与改进提供了可靠测试平台与实用评分工具。
原文摘要 · Abstract (English)
Multimodal large models are increasingly used to generate scalable vector graphics (SVG), but reliable evaluation remains underexplored. Existing protocols are often code-centric or borrow raster-image metrics after rendering SVGs, which fail to reflect human perception and overlook SVG-specific qualities such as geometry and spatial composition. We introduce SVGEval, a vision-grounded multimodal benchmark for human-aligned SVG quality assessment. SVGEval explicitly incorporates visual renderings to evaluate whether models can judge the rendered outcome rather than only inspect SVG code, and provides high-quality annotations obtained via multi-round human labeling with expert refinement. Systematic evaluations across representative multimodal models reveal a clear gap: models perform relatively well on semantic alignment and aesthetics, yet struggle on geometry- and layout-related judgments. Building on SVGEval, we train an explainable SVG quality scorer that outputs multi-aspect scores with textual rationales. Ablations show that explicit visual grounding and reasoning supervision are crucial, especially for spatial and geometric assessment. SVGEval offers a reliable testbed and practical scorer for evaluating and improving SVG generation in the era of multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。