arXiv:2504.07960cs.CV2025-04ICCV被引 41

用视觉示范实现通用图像生成,支持多种任务统一处理。

VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning

  • 通过视觉示范识别任务,避免语言指令模糊性。
  • 构建Graph200K数据集提升任务密度与知识迁移能力。
  • 复用预训练填充模型的生成先验,无需修改架构。

扩散模型的进展显著推动了各类图像生成任务的发展。然而,当前主流方法仍依赖于特定任务的模型,难以高效支持多样需求。尽管通用模型试图解决此问题,却面临任务指令泛化、任务分布合理性和架构统一等挑战。为此,我们提出VisualCloze,一个支持域内多种任务、可泛化至未见任务、统一处理多个未见任务并实现逆向生成的通用图像生成框架。不同于依赖语言指令导致任务歧义和泛化弱的方法,我们引入视觉上下文学习,使模型能从视觉示范中识别任务。同时,视觉任务分布固有的稀疏性阻碍了跨任务可迁移知识的学习。为此,我们构建Graph200K图结构数据集,建立多种相关任务,增强任务密度与可迁移知识。此外,我们发现统一图像生成形式与图像填充共享一致目标,从而可直接利用预训练填充模型的强大生成先验,无需修改架构。

原文摘要 · Abstract (English)

Recent progress in diffusion models significantly advances various image generation tasks. However, the current mainstream approach remains focused on building task-specific models, which have limited efficiency when supporting a wide range of different needs. While universal models attempt to address this limitation, they face critical challenges, including generalizable task instruction, appropriate task distributions, and unified architectural design. To tackle these challenges, we propose VisualCloze, a universal image generation framework, which supports a wide range of in-domain tasks, generalization to unseen ones, unseen unification of multiple tasks, and reverse generation. Unlike existing methods that rely on language-based task instruction, leading to task ambiguity and weak generalization, we integrate visual in-context learning, allowing models to identify tasks from visual demonstrations. Meanwhile, the inherent sparsity of visual task distributions hampers the learning of transferable knowledge across tasks. To this end, we introduce Graph200K, a graph-structured dataset that establishes various interrelated tasks, enhancing task density and transferable knowledge. Furthermore, we uncover that our unified image generation formulation shared a consistent objective with image infilling, enabling us to leverage the strong generative priors of pre-trained infilling models without modifying the architectures.

图像生成视觉上下文学习通用模型扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。