arXiv:2411.08930cs.CVcs.GR2024-11被引 2

用扩散模型精准扩展手绘图案,保持结构细节。

Structured Pattern Expansion with Diffusion Models

  • 基于扩散模型,通过局部手绘图生成大尺寸连贯图案。
  • 在多个数据集上生成的图案多样性与一致性优于现有方法。
  • 适合需要精细控制设计纹理的设计师和游戏开发人员。

最近的扩散模型在材料、纹理和3D形状合成方面取得显著进展。通过文本或图像条件化,用户可引导生成过程,大幅缩短数字资产创建时间。本文聚焦于结构化静态图案的合成问题,该领域中扩散模型通常可靠性低且难以控制。我们的方法针对图案域专门优化了扩散模型的生成能力,使用户能直接通过扩展部分手绘图案生成更大设计,同时保留输入的结构与细节。为提升图案质量,我们在图像预训练扩散模型上使用低秩适配(LoRA)进行微调,引入噪声滚动技术以确保可平铺性,并采用基于块的生成策略来支持大规模资产生成。通过全面实验验证,所提方法在生成多样化、一致性强且响应用户输入的图案方面优于现有模型。

原文摘要 · Abstract (English)

Recent advances in diffusion models have significantly improved the synthesis of materials, textures, and 3D shapes. By conditioning these models via text or images, users can guide the generation, reducing the time required to create digital assets. In this paper, we address the synthesis of structured, stationary patterns, where diffusion models are generally less reliable and, more importantly, less controllable. Our approach leverages the generative capabilities of diffusion models specifically adapted for the pattern domain. It enables users to exercise direct control over the synthesis by expanding a partially hand-drawn pattern into a larger design while preserving the structure and details of the input. To enhance pattern quality, we fine-tune an image-pretrained diffusion model on structured patterns using Low-Rank Adaptation (LoRA), apply a noise rolling technique to ensure tileability, and utilize a patch-based approach to facilitate the generation of large-scale assets. We demonstrate the effectiveness of our method through a comprehensive set of experiments, showing that it outperforms existing models in generating diverse, consistent patterns that respond directly to user input.

扩散模型图案生成可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。