arXiv:2509.07127cs.GRcs.AI2025-09中稿 · 23rd edition of In…被引 2

首个面向文本生成矢量图的真人评价对齐评估工具,兼顾视觉与语义一致性。

SVGauge: Towards Human-Aligned Evaluation for SVG Generation

  • 结合SigLIP图像嵌入与PCA白化,评估生成图的视觉保真度。
  • 通过BLIP-2生成描述并对比原提示,在SBERT与TF-IDF空间衡量语义一致性。
  • 在SHE基准上与人工评分相关性最高,适合评估零样本矢量图生成模型。

生成式可缩放矢量图形(SVG)需要针对其符号性和矢量性设计评估标准,而现有指标如FID、LPIPS或CLIPScore无法满足需求。本文提出SVGauge,首个面向文本到SVG生成的人类对齐参考评估指标。SVGauge联合测量:(i) 视觉保真度,通过提取SigLIP图像嵌入,并经主成分分析(PCA)和白化处理实现领域对齐;(ii) 语义一致性,通过比较BLIP-2生成的SVG描述与原始提示,在SBERT与TF-IDF联合空间中的相似度。在新提出的SHE基准上的评估显示,SVGauge与人工评分的相关性最高,并更准确还原了八种零样本大语言模型生成器的系统级排名。结果凸显了矢量专用评估的必要性,为未来文本到SVG生成模型提供实用基准工具。

原文摘要 · Abstract (English)

Generated Scalable Vector Graphics (SVG) images demand evaluation criteria tuned to their symbolic and vectorial nature: criteria that existing metrics such as FID, LPIPS, or CLIPScore fail to satisfy. In this paper, we introduce SVGauge, the first human-aligned, reference based metric for text-to-SVG generation. SVGauge jointly measures (i) visual fidelity, obtained by extracting SigLIP image embeddings and refining them with PCA and whitening for domain alignment, and (ii) semantic consistency, captured by comparing BLIP-2-generated captions of the SVGs against the original prompts in the combined space of SBERT and TF-IDF. Evaluation on the proposed SHE benchmark shows that SVGauge attains the highest correlation with human judgments and reproduces system-level rankings of eight zero-shot LLM-based generators more faithfully than existing metrics. Our results highlight the necessity of vector-specific evaluation and provide a practical tool for benchmarking future text-to-SVG generation models.

矢量图生成评估指标语义一致人类对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。