让模型从真实图像中可靠生成SVG,解决噪声与杂乱干扰问题
WildSVG: Towards Reliable SVG Generation Under Real-Word Conditions
- 构建自然与合成双数据集,模拟真实世界复杂场景
- 现有模型在真实图像上生成效果显著下降,性能远未达标
- 迭代优化方法展现潜力,适合做视觉矢量生成研究者参考
我们提出SVG提取任务,旨在将图像中的特定视觉内容转化为可缩放矢量图形。现有多模态模型在干净渲染图或文本描述下表现良好,但在包含噪声、杂乱和域偏移的真实图像中表现不佳。核心挑战在于缺乏合适基准。为此,我们引入WildSVG基准,包含两个互补数据集:自然野SVG(基于真实公司标志图像及其对应的SVG标注)和合成野SVG(将复杂SVG渲染嵌入真实场景以模拟困难条件)。二者共同构成首个系统性评估SVG提取的基准。我们对当前先进多模态模型进行测试,发现其在真实场景下的表现远低于可靠应用所需水平。然而,迭代精炼方法展现出前景,模型能力正稳步提升。
原文摘要 · Abstract (English)
We introduce the task of SVG extraction, which consists in translating specific visual inputs from an image into scalable vector graphics. Existing multimodal models achieve strong results when generating SVGs from clean renderings or textual descriptions, but they fall short in real-world scenarios where natural images introduce noise, clutter, and domain shifts. A central challenge in this direction is the lack of suitable benchmarks. To address this need, we introduce the WildSVG Benchmark, formed by two complementary datasets: Natural WildSVG, built from real images containing company logos paired with their SVG annotations, and Synthetic WildSVG, which blends complex SVG renderings into real scenes to simulate difficult conditions. Together, these resources provide the first foundation for systematic benchmarking SVG extraction. We benchmark state-of-the-art multimodal models and find that current approaches perform well below what is needed for reliable SVG extraction in real scenarios. Nonetheless, iterative refinement methods point to a promising path forward, and model capabilities are steadily improving
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。