MagicQuill让用户用自然语言快速实现图像编辑。
MagicQuill: An Intelligent Interactive Image Editing System
- 通过多模态大模型实时理解用户意图,无需输入提示词。
- 采用增强扩散先验与双分支模块,实现精准图像编辑。
- 适合创意设计、非专业用户快速实现视觉构思。
图像编辑涉及多种复杂任务,需要高效精确的操作手段。本文提出MagicQuill,一个集成的交互式图像编辑系统,可快速实现创意构想。系统具备简洁但功能强大的界面,仅需少量输入即可完成插入元素、删除物体、改变颜色等操作。这些交互由多模态大语言模型(MLLM)实时监控,自动预测编辑意图,无需显式输入提示。最后,利用经过精心设计的双分支插件模块增强的扩散先验,对编辑请求进行高精度处理。实验结果表明,MagicQuill在高质量图像编辑方面表现优异。请访问 https://magic-quill.github.io 体验系统。
原文摘要 · Abstract (English)
Image editing involves a variety of complex tasks and requires efficient and precise manipulation techniques. In this paper, we present MagicQuill, an integrated image editing system that enables swift actualization of creative ideas. Our system features a streamlined yet functionally robust interface, allowing for the articulation of editing operations (e.g., inserting elements, erasing objects, altering color) with minimal input. These interactions are monitored by a multimodal large language model (MLLM) to anticipate editing intentions in real time, bypassing the need for explicit prompt entry. Finally, we apply a powerful diffusion prior, enhanced by a carefully learned two-branch plug-in module, to process editing requests with precise control. Experimental results demonstrate the effectiveness of MagicQuill in achieving high-quality image edits. Please visit https://magic-quill.github.io to try out our system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。