用少量样本教会扩散模型理解并迁移视觉编辑意图。
RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers
- 通过源图-目标图对提取内容感知的编辑意图。
- 在218类任务上显著提升生成质量和编辑效果。
- 适合需要少样本视觉编辑的AI绘画与图像处理场景
受大语言模型上下文学习机制启发,基于视觉提示的通用图像编辑新范式正兴起。现有单参考方法多聚焦风格或外观调整,难以应对非刚性变换。为此,我们提出利用源-目标图像对提取并迁移内容感知的编辑意图至新查询图像。为此,引入轻量级模块RelationAdapter,使基于扩散变压器(DiT)的模型能有效捕捉和应用最小样本下的视觉变换。同时构建了包含218种多样化编辑任务的Relation252K数据集,用于评估模型在视觉提示驱动场景下的泛化与适应能力。在Relation252K上的实验表明,RelationAdapter显著提升了模型理解与传递编辑意图的能力,生成质量与整体编辑性能均有明显提升。
原文摘要 · Abstract (English)
Inspired by the in-context learning mechanism of large language models (LLMs), a new paradigm of generalizable visual prompt-based image editing is emerging. Existing single-reference methods typically focus on style or appearance adjustments and struggle with non-rigid transformations. To address these limitations, we propose leveraging source-target image pairs to extract and transfer content-aware editing intent to novel query images. To this end, we introduce RelationAdapter, a lightweight module that enables Diffusion Transformer (DiT) based models to effectively capture and apply visual transformations from minimal examples. We also introduce Relation252K, a comprehensive dataset comprising 218 diverse editing tasks, to evaluate model generalization and adaptability in visual prompt-driven scenarios. Experiments on Relation252K show that RelationAdapter significantly improves the model's ability to understand and transfer editing intent, leading to notable gains in generation quality and overall editing performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。