arXiv:2502.14007cs.GRcs.AI2025-02中稿 · The International …

用预训练扩散模型生成高保真手绘草图转图像,无需重训练。

d-Sketch: Improving Visual Fidelity of Sketch-to-Image Translation with Pretrained Latent Diffusion Models without Retraining

  • 通过轻量映射网络在潜在空间转换草图特征
  • 在高分辨率下实现逼真图像生成,优于现有方法
  • 适合内容创作与快速原型设计,无需数据重训练

图像到图像的结构化引导可对合成图像的形状实现精细控制。从用户指定的粗略手绘草图生成高质量真实感图像是一项任务,旨在对条件生成过程施加结构约束。尽管该任务在内容创作和学术研究中具有广泛价值,但自由手绘草图中的显著模糊性使其本质困难。此外,在形状一致性与真实感生成之间的权衡也增加了复杂性。基于生成对抗网络(GAN)的现有方法通常依赖条件GAN或GAN反演,常需特定应用的数据与优化目标。最近提出的去噪扩散概率模型(DDPM)在一般图像合成的低层视觉属性上实现质的飞跃。然而,直接在特定子任务上对大规模扩散模型进行重训练往往因计算成本高昂和数据不足而难以实现。本文提出一种无需重训练即可利用大规模扩散模型特性的草图转图像技术。具体而言,我们使用可学习的轻量级映射网络实现源域到目标域的潜在特征转换。实验表明,所提方法在定性和定量基准上均优于现有技术,能够从粗糙手绘草图生成高分辨率真实感图像。

原文摘要 · Abstract (English)

Structural guidance in an image-to-image translation allows intricate control over the shapes of synthesized images. Generating high-quality realistic images from user-specified rough hand-drawn sketches is one such task that aims to impose a structural constraint on the conditional generation process. While the premise is intriguing for numerous use cases of content creation and academic research, the problem becomes fundamentally challenging due to substantial ambiguities in freehand sketches. Furthermore, balancing the trade-off between shape consistency and realistic generation contributes to additional complexity in the process. Existing approaches based on Generative Adversarial Networks (GANs) generally utilize conditional GANs or GAN inversions, often requiring application-specific data and optimization objectives. The recent introduction of Denoising Diffusion Probabilistic Models (DDPMs) achieves a generational leap for low-level visual attributes in general image synthesis. However, directly retraining a large-scale diffusion model on a domain-specific subtask is often extremely difficult due to demanding computation costs and insufficient data. In this paper, we introduce a technique for sketch-to-image translation by exploiting the feature generalization capabilities of a large-scale diffusion model without retraining. In particular, we use a learnable lightweight mapping network to achieve latent feature translation from source to target domain. Experimental results demonstrate that the proposed method outperforms the existing techniques in qualitative and quantitative benchmarks, allowing high-resolution realistic image synthesis from rough hand-drawn sketches.

草图生成扩散模型图像合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。