arXiv:2502.14397cs.CV2025-02被引 27

用少量样例学习艺术家风格,实现自然融合的图像涂鸦。

PhotoDoodle: Learning Artistic Image Editing from Few-Shot Pairwise Data

  • 两阶段训练:先通用模型,再小样本微调捕捉独特风格。
  • 支持无缝融合、透视对齐与背景无畸变的高质量图像涂鸦。
  • 适合数字艺术创作,尤其擅长个性化风格迁移。

我们提出PhotoDoodle,一种新型图像编辑框架,旨在支持艺术家在照片上叠加装饰元素进行图像涂鸦。该任务难点在于插入元素需与背景真实融合,实现透视对齐和上下文一致,同时保持背景无变形,并高效捕捉艺术家的独特风格。现有方法多聚焦全局风格迁移或局部修复,难以满足上述需求。为此,PhotoDoodle采用两阶段训练策略:首先使用大规模数据训练通用编辑模型OmniEditor;随后利用艺术家精心标注的小规模前后图像对数据集EditLoRA对模型进行微调,以捕获具体编辑风格与技巧。为提升生成结果的一致性,引入位置编码复用机制。此外,我们发布了包含六种高质量风格的PhotoDoodle数据集。大量实验表明,该方法在定制化图像编辑中表现卓越且鲁棒,为艺术创作开辟了新可能。

原文摘要 · Abstract (English)

We introduce PhotoDoodle, a novel image editing framework designed to facilitate photo doodling by enabling artists to overlay decorative elements onto photographs. Photo doodling is challenging because the inserted elements must appear seamlessly integrated with the background, requiring realistic blending, perspective alignment, and contextual coherence. Additionally, the background must be preserved without distortion, and the artist's unique style must be captured efficiently from limited training data. These requirements are not addressed by previous methods that primarily focus on global style transfer or regional inpainting. The proposed method, PhotoDoodle, employs a two-stage training strategy. Initially, we train a general-purpose image editing model, OmniEditor, using large-scale data. Subsequently, we fine-tune this model with EditLoRA using a small, artist-curated dataset of before-and-after image pairs to capture distinct editing styles and techniques. To enhance consistency in the generated results, we introduce a positional encoding reuse mechanism. Additionally, we release a PhotoDoodle dataset featuring six high-quality styles. Extensive experiments demonstrate the advanced performance and robustness of our method in customized image editing, opening new possibilities for artistic creation.

图像编辑风格迁移小样本学习艺术创作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。