arXiv:2505.20271cs.CVcs.AI2025-05SIGGRAPH被引 8

无需训练,用提示词精准把物体插入图片,效果更自然。

In-Context Brush: Zero-shot Customized Subject Insertion with Context-Aware Latent Space Manipulation

  • 用文本和图像当提示,动态调整注意力机制实现零样本插入
  • 在注意力头内和头间双重操控潜空间,提升对提示的响应能力
  • 适合需要快速定制化图像生成的设计师或内容创作者

扩散模型在多模态引导图像生成方面取得进展,实现了用户指定物体的无缝插入。但现有方法在保持高保真度和准确对齐文本意图方面仍存在挑战。本文提出「In-Context Brush」,一种基于上下文学习范式的零样本定制物体插入框架。将物体图像与文本提示视为跨模态演示,目标图像中掩码区域作为查询,旨在不微调模型的前提下,根据提示完成高质量修复。基于预训练的MMDiT inpainting网络,在测试时通过双层潜空间操控实现增强:在注意力头内部进行“潜特征偏移”,动态调整关注内容以反映目标语义;在头间进行“注意力重加权”,通过差异化优先级放大提示控制力。大量实验与应用表明,本方法在身份保留、文本对齐与图像质量上均优于现有最先进方法,且无需专门训练或额外数据收集。

原文摘要 · Abstract (English)

Recent advances in diffusion models have enhanced multimodal-guided visual generation, enabling customized subject insertion that seamlessly "brushes" user-specified objects into a given image guided by textual prompts. However, existing methods often struggle to insert customized subjects with high fidelity and align results with the user's intent through textual prompts. In this work, we propose "In-Context Brush", a zero-shot framework for customized subject insertion by reformulating the task within the paradigm of in-context learning. Without loss of generality, we formulate the object image and the textual prompts as cross-modal demonstrations, and the target image with the masked region as the query. The goal is to inpaint the target image with the subject aligning textual prompts without model tuning. Building upon a pretrained MMDiT-based inpainting network, we perform test-time enhancement via dual-level latent space manipulation: intra-head "latent feature shifting" within each attention head that dynamically shifts attention outputs to reflect the desired subject semantics and inter-head "attention reweighting" across different heads that amplifies prompt controllability through differential attention prioritization. Extensive experiments and applications demonstrate that our approach achieves superior identity preservation, text alignment, and image quality compared to existing state-of-the-art methods, without requiring dedicated training or additional data collection.

图像生成零样本扩散模型提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。