arXiv:2605.24398cs.CVcs.AI2026-05

用圆角多边形表示图像矢量化,提升真实场景下的效果与鲁棒性。

VectorArk: Learning Practical Image Vectorization with Rounded Polygon Representation

论文配图:VectorArk: Learning Practical Image Vectorization with Rounded Polygon Representation
图 1 · 摘自论文原文
  • 采用圆角多边形表示法,简化学习并生成平滑美观的矢量图
  • 在多个数据集上实现更高几何完整性和更低伪影,优于现有方法
  • 适合处理真实图像、文本生成图像等复杂输入场景

基于视觉语言模型(VLM)的图像矢量化方法虽在合成基准上表现优异,但普遍仅在高分辨率位图重矢量化任务中评估,难以泛化至真实场景——如未知位图化方法或由文生图模型生成的图像。本文提出VectorArk,一种面向实际应用的新型VLM矢量化模型。其采用新颖的圆角多边形表示,降低学习难度的同时自然生成光滑、视觉优美的矢量基元。同时设计退化模型以增强对多样化、不完美输入的鲁棒性。实验表明,相较于以往方法,VectorArk在多个数据集上均实现更优的几何完整性与伪影抑制效果;全面消融实验验证了各组件的有效性。

原文摘要 · Abstract (English)

Recent vision-language model (VLM)-based approaches have achieved impressive results on image vectorization tasks. However, they are typically evaluated on synthetic benchmarks, where clean SVGs are rasterized at high resolution and then re-vectorized. As a result, these methods generalize poorly to real-world scenarios, such as images with unknown rasterization methods or those generated by text-to-image models. We introduce VectorArk, a new VLM-based model designed for robust and practical image vectorization. VectorArk employs a novel rounded polygon representation that simplifies the learning process while naturally producing smooth, visually appealing primitives. We also propose a degradation model that enhances robustness across diverse and imperfect inputs. Our experiments show that, in contrast to previous methods, VectorArk achieves superior geometric completeness and artifact suppression across multiple datasets, with comprehensive ablations validating the contribution of each component.

图像矢量化视觉语言模型圆角多边形鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。