arXiv:2412.01027cs.CV2024-12CVPR被引 14

让自回归模型通过少量示例快速学会文本+图像指令的图像操作。

Unleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulation

  • 用分阶段注意力机制拆解学习与应用过程,提升推理能力。
  • 在人类评估中比现有方法高出至少19%。
  • 适合需要快速适配新图像操作任务的开发者使用。

近年来,文本引导的图像操作取得了显著进展。为缓解语言描述的歧义性,少样本学习结合视觉示例被用于训练数据中罕见或难以纯语言描述的指令。然而,从视觉提示中学习需要强大的推理能力,而扩散模型在此方面表现不足。为此,我们提出一种新型多模态自回归模型InstaManip,可通过上下文学习即时从文本和视觉指导中学习新的图像操作,并应用于新查询图像。具体地,我们设计了一种创新的组自注意力机制,将上下文学习过程分为学习与应用两个阶段,从而简化复杂问题。同时引入关系正则化方法,进一步分离出图像变换特征与示例图像中的无关内容。大量实验表明,该方法在人类评估中显著优于先前的少样本图像操作模型(提升≥19%)。此外,增加示例图像的数量或多样性可进一步提升性能。

原文摘要 · Abstract (English)

Text-guided image manipulation has experienced notable advancement in recent years. In order to mitigate linguistic ambiguity, few-shot learning with visual examples has been applied for instructions that are underrepresented in the training set, or difficult to describe purely in language. However, learning from visual prompts requires strong reasoning capability, which diffusion models are struggling with. To address this issue, we introduce a novel multi-modal autoregressive model, dubbed $\textbf{InstaManip}$, that can $\textbf{insta}$ntly learn a new image $\textbf{manip}$ulation operation from textual and visual guidance via in-context learning, and apply it to new query images. Specifically, we propose an innovative group self-attention mechanism to break down the in-context learning process into two separate stages -- learning and applying, which simplifies the complex problem into two easier tasks. We also introduce a relation regularization method to further disentangle image transformation features from irrelevant contents in exemplar images. Extensive experiments suggest that our method surpasses previous few-shot image manipulation models by a notable margin ($\geq$19% in human evaluation). We also find our model can be further boosted by increasing the number or diversity of exemplar images.

图像操作少样本学习上下文学习自回归模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。