用迭代优化让通用视觉模型生成矢量图,发现编解码改进有效但自我修正能力不足。
Evaluating Constrained Iterative Refinement for Scalable Vector Graphics Generation with Off-the-Shelf VLMs

- 结合视觉反馈与结构化编辑的迭代优化方法
- 约束解码提升编译成功率,最高达87%
- 适合研究多模态生成与可扩展图形技术的开发者
可缩放矢量图形(SVG)支撑现代视觉生态,但当前主流生成模型集中于位图图像。本文探究推理阶段方法是否能激活现成视觉语言模型(VLM)的SVG生成能力。通过系统评估一种结合视觉反馈、结构化编辑和约束解码的受限迭代优化框架,我们发现约束解码显著提升编译成功率,而迭代优化暴露出现有VLM在视觉推理与自我修正方面的缺陷。实验覆盖多个VLM及生成设置,结果揭示了利用推理时方法适配通用VLM进行SVG生成的潜力与当前局限。
原文摘要 · Abstract (English)
Scalable Vector Graphics (SVGs) power much of the modern visual ecosystem, yet state-of-the-art generative models focus almost entirely on rasterized images. We explore whether inference-time methods can unlock SVG generation capabilities in off-the-shelf vision-language models (VLMs). We systematically evaluate a constrained iterative refinement harness that combines visual feedback, structured editing, and constrained decoding to characterize the capabilities and limitations of current VLMs for SVG generation. Across multiple VLMs and generation settings, we find that constrained decoding improves compilation success rates, while iterative refinement reveals a deficit in visual reasoning and self-correction. Our results highlight both the promise and current limitations of using inference-time methods to adapt general-purpose VLMs for SVG generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。