用示例图和文本实现高效精准的图像编辑,无需额外训练
ReEdit: Multimodal Exemplar-Based Image Editing with Diffusion Models
- 通过图文双模态捕捉示例对的编辑动作,实现端到端编辑
- 在多个数据集上优于现有方法,且速度比最优基线快4倍
- 无需微调,适合快速实际应用,尤其适合非专业用户
现代文本到图像(T2I)扩散模型推动了高质量逼真图像的生成。尽管当前主流编辑方式依赖文本指令,但自然语言与图像间复杂的多对多映射关系使其难以精确控制。本文针对示例式图像编辑任务——即从示例对中迁移编辑效果至内容图像——提出ReEdit框架。该框架为模块化、高效的端到端设计,同时捕捉文本与图像模态中的编辑信息,并保证编辑后图像的保真度。通过与前沿基线的广泛对比及关键设计选择的敏感性分析,结果表明ReEdit在定性和定量指标上均持续领先。此外,该方法具有高实用性:无需任务特定优化,运行速度为最优基线的4倍。
原文摘要 · Abstract (English)
Modern Text-to-Image (T2I) Diffusion models have revolutionized image editing by enabling the generation of high-quality photorealistic images. While the de facto method for performing edits with T2I models is through text instructions, this approach non-trivial due to the complex many-to-many mapping between natural language and images. In this work, we address exemplar-based image editing -- the task of transferring an edit from an exemplar pair to a content image(s). We propose ReEdit, a modular and efficient end-to-end framework that captures edits in both text and image modalities while ensuring the fidelity of the edited image. We validate the effectiveness of ReEdit through extensive comparisons with state-of-the-art baselines and sensitivity analyses of key design choices. Our results demonstrate that ReEdit consistently outperforms contemporary approaches both qualitatively and quantitatively. Additionally, ReEdit boasts high practical applicability, as it does not require any task-specific optimization and is four times faster than the next best baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。