KnobGen让草图生成图像更灵活,新手和高手都能用。
KnobGen: Controlling the Sophistication of Artwork in Sketch-Based Diffusion Models
- 双路径设计:粗粒度控制器抓整体,细粒度控制器调细节。
- 可调节控制旋钮,根据草图复杂度自动适配输出精度。
- 适合想轻松作画的用户或需要精细控制的专业人士。
扩散模型在文本到图像生成上取得显著进展,但难以兼顾细节精度与高层控制。现有方法如ControlNet和T2I-Adapter虽能精准跟随专业草图,却对新手草图中的瑕疵过度敏感;而粗粒度抽象框架虽易用,又缺乏精细控制能力。为此,我们提出KnobGen,一种双路径框架,可无缝适应不同复杂度草图与用户技能水平。其包含粗粒度控制器(CGC)捕捉高层语义,细粒度控制器(FGC)实现细节优化,通过可控旋钮机制调节两者权重,灵活响应输入质量。在MultiGen-20M数据集及新收集的草图数据集上验证,该方法既能保持图像自然外观,又能有效控制输出结果。
原文摘要 · Abstract (English)
Recent advances in diffusion models have significantly improved text-to-image (T2I) generation, but they often struggle to balance fine-grained precision with high-level control. Methods like ControlNet and T2I-Adapter excel at following sketches by seasoned artists but tend to be overly rigid, replicating unintentional flaws in sketches from novice users. Meanwhile, coarse-grained methods, such as sketch-based abstraction frameworks, offer more accessible input handling but lack the precise control needed for detailed, professional use. To address these limitations, we propose KnobGen, a dual-pathway framework that democratizes sketch-based image generation by seamlessly adapting to varying levels of sketch complexity and user skill. KnobGen uses a Coarse-Grained Controller (CGC) module for high-level semantics and a Fine-Grained Controller (FGC) module for detailed refinement. The relative strength of these two modules can be adjusted through our knob inference mechanism to align with the user's specific needs. These mechanisms ensure that KnobGen can flexibly generate images from both novice sketches and those drawn by seasoned artists. This maintains control over the final output while preserving the natural appearance of the image, as evidenced on the MultiGen-20M dataset and a newly collected sketch dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。