用视觉问答反馈提升扩散模型画手绘素描的精准度
StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback
- 用强化学习+视觉问答优化生成过程,增强文本对齐
- 在提示词一致性和风格保真度上优于稳定扩散基线
- 首次发布含问答对的手绘素描数据集SketchDUO
尽管扩散模型在图像生成质量上取得显著进展,但在生成像素级手绘素描(一种抽象表达的代表)方面仍面临挑战。为此,我们提出StableSketcher,一种新框架,使扩散模型能生成高提示保真度的手绘素描。该框架通过微调变分自编码器优化潜在空间解码,更准确捕捉素描特征;同时引入基于视觉问答的奖励函数,用于强化学习,提升文本-图像对齐与语义一致性。大量实验表明,StableSketcher生成的素描在风格保真度和提示匹配度上均优于稳定扩散基线。此外,我们提出了SketchDUO——目前已知首个包含实例级素描、配图文本及问答对的数据集,弥补了现有数据集依赖图像-标签对的局限。代码与数据集将在论文接受后公开。项目页:https://zihos.github.io/StableSketcher
原文摘要 · Abstract (English)
Although recent advancements in diffusion models have significantly enriched the quality of generated images, challenges remain in synthesizing pixel-based human-drawn sketches, a representative example of abstract expression. To combat these challenges, we propose StableSketcher, a novel framework that empowers diffusion models to generate hand-drawn sketches with high prompt fidelity. Within this framework, we fine-tune the variational autoencoder to optimize latent decoding, enabling it to better capture the characteristics of sketches. In parallel, we integrate a new reward function for reinforcement learning based on visual question answering, which improves text-image alignment and semantic consistency. Extensive experiments demonstrate that StableSketcher generates sketches with improved stylistic fidelity, achieving better alignment with prompts compared to the Stable Diffusion baseline. Additionally, we introduce SketchDUO, to the best of our knowledge, the first dataset comprising instance-level sketches paired with captions and question-answer pairs, thereby addressing the limitations of existing datasets that rely on image-label pairs. Our code and dataset will be made publicly available upon acceptance. Project page: https://zihos.github.io/StableSketcher
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。