用视觉信号引导生成,让SVG更准确美观
DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance
- 端到端联合生成图像和SVG tokens
- 测试时利用模型自身视觉预测提升生成质量
- 适合需要高保真、结构清晰SVG的场景
基于视觉语言模型的SVG生成方法虽已取得显著进展,但因解码过程仅依赖文本而缺乏视觉信号,常在复杂语义下表现不佳,难以生成视觉上逼真或几何上一致的SVG。我们提出DuetSVG,一种统一的多模态模型,以端到端方式联合生成图像标记和对应SVG标记。该模型在图像与SVG数据集上共同训练。推理时,采用一种新颖的测试时缩放策略,利用模型原生的视觉预测作为引导,提升SVG解码质量。大量实验表明,本方法在多种应用场景中均优于现有方法,生成的SVG具备高度视觉真实性、语义一致性与语法整洁性。
原文摘要 · Abstract (English)
Recent vision-language model (VLM)-based approaches have achieved impressive results on SVG generation. However, because they generate only text and lack visual signals during decoding, they often struggle with complex semantics and fail to produce visually appealing or geometrically coherent SVGs. We introduce DuetSVG, a unified multimodal model that jointly generates image tokens and corresponding SVG tokens in an end-to-end manner. DuetSVG is trained on both image and SVG datasets. At inference, we apply a novel test-time scaling strategy that leverages the model's native visual predictions as guidance to improve SVG decoding quality. Extensive experiments show that our method outperforms existing methods, producing visually faithful, semantically aligned, and syntactically clean SVGs across a wide range of applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。