arXiv:2605.01135cs.CV2026-05

用手绘草图+文字生成图像编辑数据集,提升精准控制能力

ScribbleEdit: Synthetic Data for Image Editing with Scribbles and Text

论文配图:ScribbleEdit: Synthetic Data for Image Editing with Scribbles and Text
图 1 · 摘自论文原文
  • 自动生成含草图与文本的图像编辑数据对
  • 微调后模型在空间对齐与语义一致上显著提升
  • 适合需要精细图文联合控制的编辑任务

生成模型在图像编辑方面取得显著进展,但用户仍难以实现精确且直观的控制。自然语言可表达纹理、颜色等语义信息,但缺乏空间精度;手绘草图能提供粗略边界,却无法描述具体视觉属性。因此,精确编辑需结合两种模态。然而,现有模型缺乏针对抽象草图与文本联合输入的训练数据,难以有效理解。为此,我们提出ScribbleEdit,一个大规模合成数据集,通过自动修复(inpainting)生成源-目标图像对,并配以人工绘制的草图和基于视觉语言模型(VLM)生成的文本指令。利用该数据集,我们评估并微调了基于扩散和自回归的统一多模态图像编辑模型。实验表明,预训练模型对抽象草图表现不佳,但在ScribbleEdit上微调后,生成结果的空间对齐性与语义一致性均明显改善。

原文摘要 · Abstract (English)

Recent progress in generative models has significantly advanced image editing capabilities, yet precise and intuitive user control remains difficult. Specifically, users often struggle to communicate both exact spatial layouts and specific semantic details simultaneously. While natural language instructions effectively convey high-level semantics like texture and color, they lack spatial specificity. Conversely, freehand scribbles provide rough spatial boundaries but cannot express detailed visual attributes. Consequently, achieving precise control requires combining both modalities. However, existing models struggle to jointly interpret abstract scribbles alongside text due to a lack of specialized training data. In this work, we introduce ScribbleEdit, a large-scale synthetic dataset designed to bridge this gap by combining natural language instructions with freehand scribble inputs for more accurate, controllable edits. We construct this dataset through a synthetic pipeline that automatically generates source-target image pairs via inpainting, which are then paired with human-drawn scribbles and VLM-generated text instructions. Using ScribbleEdit, we evaluate and finetune both diffusion-based and autoregressive unified multimodal image editing models. Our experiments reveal that while off-the-shelf models struggle with abstract scribble inputs, finetuning on our synthetic dataset significantly improves their ability to generate spatially aligned and semantically consistent edits.

图像编辑草图控制合成数据多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。