用扩散模型1秒生成高质量矢量草图,速度快且风格自然。
SwiftSketch: A Diffusion Model for Image-to-Vector Sketch Generation
- 基于扩散模型逐步去噪控制点,直接生成矢量草图。
- 生成时间少于1秒,且保持高保真与艺术感。
- 适合需要快速出图的UI设计、创意工具场景。
近期大型视觉语言模型已实现高度表达且多样的矢量草图生成,但现有方法依赖耗时的优化过程,需反复调用预训练模型确定笔画位置,限制了实际应用。本文提出SwiftSketch,一种图像条件下的矢量草图生成扩散模型,可在1秒内生成高质量草图。该模型通过从高斯分布中采样并逐步去噪笔画控制点来工作,采用Transformer解码器架构,有效处理矢量表示的离散性,并捕捉笔画间的全局依赖关系。为训练模型,构建了合成图像-草图配对数据集,克服了现有数据集由非艺术家创作、专业度不足的问题。合成草图使用ControlSketch生成,通过引入深度感知的ControlNet增强基于SDS的技术,实现精确空间控制。实验表明,SwiftSketch在多样化概念上具有良好泛化能力,高效生成兼具高保真度与自然美观风格的草图。
原文摘要 · Abstract (English)
Recent advancements in large vision-language models have enabled highly expressive and diverse vector sketch generation. However, state-of-the-art methods rely on a time-consuming optimization process involving repeated feedback from a pretrained model to determine stroke placement. Consequently, despite producing impressive sketches, these methods are limited in practical applications. In this work, we introduce SwiftSketch, a diffusion model for image-conditioned vector sketch generation that can produce high-quality sketches in less than a second. SwiftSketch operates by progressively denoising stroke control points sampled from a Gaussian distribution. Its transformer-decoder architecture is designed to effectively handle the discrete nature of vector representation and capture the inherent global dependencies between strokes. To train SwiftSketch, we construct a synthetic dataset of image-sketch pairs, addressing the limitations of existing sketch datasets, which are often created by non-artists and lack professional quality. For generating these synthetic sketches, we introduce ControlSketch, a method that enhances SDS-based techniques by incorporating precise spatial control through a depth-aware ControlNet. We demonstrate that SwiftSketch generalizes across diverse concepts, efficiently producing sketches that combine high fidelity with a natural and visually appealing style.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。