无需训练即可精准插入物体,同时控制形状和风格。
FreeInsert: Personalized Object Insertion with Geometric and Style Control
- 利用3D重建实现物体的几何控制,支持视角与形状调整。
- 通过扩散适配器保持插入物与背景的风格一致。
- 无需微调模型,适合快速个性化图像编辑。
文本到图像的扩散模型在图像生成方面取得了显著进展,实现了便捷的个性化生成。然而,现有图像编辑方法在处理个性化图像构图任务时仍存在局限:一是缺乏对插入物体的几何控制,当前方法多局限于2D空间且依赖文本指令,难以精确控制物体几何;二是风格一致性挑战,现有方法常忽略插入物与背景的风格匹配,导致真实感不足;三是插入物体无需大量训练仍具挑战性。为此,我们提出全新的免训练框架FreeInsert,通过利用3D几何信息,将2D物体转换为3D,进行交互式3D编辑后,从指定视角重新渲染为2D图像。该过程引入了形状、视角等几何控制,并结合扩散适配器实现风格与内容控制,最终通过扩散模型生成几何可控、风格一致的编辑图像。
原文摘要 · Abstract (English)
Text-to-image diffusion models have made significant progress in image generation, allowing for effortless customized generation. However, existing image editing methods still face certain limitations when dealing with personalized image composition tasks. First, there is the issue of lack of geometric control over the inserted objects. Current methods are confined to 2D space and typically rely on textual instructions, making it challenging to maintain precise geometric control over the objects. Second, there is the challenge of style consistency. Existing methods often overlook the style consistency between the inserted object and the background, resulting in a lack of realism. In addition, the challenge of inserting objects into images without extensive training remains significant. To address these issues, we propose \textit{FreeInsert}, a novel training-free framework that customizes object insertion into arbitrary scenes by leveraging 3D geometric information. Benefiting from the advances in existing 3D generation models, we first convert the 2D object into 3D, perform interactive editing at the 3D level, and then re-render it into a 2D image from a specified view. This process introduces geometric controls such as shape or view. The rendered image, serving as geometric control, is combined with style and content control achieved through diffusion adapters, ultimately producing geometrically controlled, style-consistent edited images via the diffusion model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。