arXiv:2412.12463cs.CVcs.AI2024-12CVPR被引 1

用类比学习编程式图像编辑,让复杂图案修改更直观。

Pattern Analogies: Learning to Perform Programmatic Image Edits by Analogy

  • 通过一对简单图案类比,指导生成模型执行结构化编辑。
  • 在真实艺术家图案上实现精准编辑,并泛化到未见风格。
  • 自动生成高质量合成数据集,支持模型高效训练。

模式图像广泛存在于数字与物理世界,其编辑工具极具价值。但模式图像的编辑颇具挑战:目标往往是程序化编辑——即修改生成该模式的底层程序。现有方法难以推断复杂图像的生成程序,所得程序杂乱无章,导致编辑困难。本文提出一种新方法,通过使用模式类比(一对简单模式展示预期编辑)和基于学习的生成模型,实现直观的程序化编辑。为此,我们设计了专用领域语言 SplitWeave,结合合成类比采样框架,构建大规模高质量合成训练数据集。同时提出 TriFuser,一种专为解决直接应用扩散模型时关键问题而设计的潜在扩散模型。在真实世界、艺术家提供的图案上的大量实验表明,本方法能忠实执行所演示的编辑,并泛化至训练分布外的相关风格。

原文摘要 · Abstract (English)

Pattern images are everywhere in the digital and physical worlds, and tools to edit them are valuable. But editing pattern images is tricky: desired edits are often programmatic: structure-aware edits that alter the underlying program which generates the pattern. One could attempt to infer this underlying program, but current methods for doing so struggle with complex images and produce unorganized programs that make editing tedious. In this work, we introduce a novel approach to perform programmatic edits on pattern images. By using a pattern analogy -- a pair of simple patterns to demonstrate the intended edit -- and a learning-based generative model to execute these edits, our method allows users to intuitively edit patterns. To enable this paradigm, we introduce SplitWeave, a domain-specific language that, combined with a framework for sampling synthetic pattern analogies, enables the creation of a large, high-quality synthetic training dataset. We also present TriFuser, a Latent Diffusion Model (LDM) designed to overcome critical issues that arise when naively deploying LDMs to this task. Extensive experiments on real-world, artist-sourced patterns reveals that our method faithfully performs the demonstrated edit while also generalizing to related pattern styles beyond its training distribution.

图像编辑程序化生成类比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。