X-Prompt让视觉语言模型通过上下文示例实现通用图像生成。
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
- 纯自回归架构,用上下文示例压缩特征支持长序列输入。
- 在已见和未见任务上均达到竞争力性能,支持跨任务泛化。
- 适合需要灵活生成多种图像的开发者与研究者使用。
上下文生成是大型语言模型(LLMs)开放任务泛化能力的关键。通过少量示例作为上下文,LLMs 可完成领域内和领域外任务。基于 LLMs 构建的自回归视觉语言模型(VLMs)在文本到图像生成方面表现出色。然而,上下文学习在通用图像生成中的潜力仍待探索。为此,我们提出 X-Prompt,一种纯自回归的大规模视觉语言模型,可在统一的上下文学习框架下,在多种已见和未见图像生成任务中实现具有竞争力的表现。X-Prompt 采用专用设计,高效压缩上下文示例中的关键特征,支持更长的上下文序列,提升对未见任务的泛化能力。统一的文本与图像预测训练任务使 X-Prompt 能够增强对上下文示例的任务感知,从而更好地处理通用图像生成。大量实验验证了其在多样化已见任务上的表现以及对未见任务的泛化能力。
原文摘要 · Abstract (English)
In-context generation is a key component of large language models' (LLMs) open-task generalization capability. By leveraging a few examples as context, LLMs can perform both in-domain and out-of-domain tasks. Recent advancements in auto-regressive vision-language models (VLMs) built upon LLMs have showcased impressive performance in text-to-image generation. However, the potential of in-context learning for general image generation tasks remains largely unexplored. To address this, we introduce X-Prompt, a purely auto-regressive large-vision language model designed to deliver competitive performance across a wide range of both seen and unseen image generation tasks, all within a unified in-context learning framework. X-Prompt incorporates a specialized design that efficiently compresses valuable features from in-context examples, supporting longer in-context token sequences and improving its ability to generalize to unseen tasks. A unified training task for both text and image prediction enables X-Prompt to handle general image generation with enhanced task awareness from in-context examples. Extensive experiments validate the model's performance across diverse seen image generation tasks and its capacity to generalize to previously unseen tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。