用大模型的对偶生成特性实现零样本主体图像生成,精准还原主体特征。
Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator
- 将主体生成重构为左右图对的修复任务,通过文本条件补全右图。
- 在用户评估中显著优于传统零样本方法,主体对齐更准确。
- 支持主体生成、风格迁移与编辑,适用于多种图像生成场景。
主体驱动的文本到图像生成旨在根据文本提示生成特定主体的新图像,需同时捕捉主体视觉特征与文本语义内容。传统方法依赖耗时耗资源的微调进行主体对齐,而近期零样本方法采用实时图像提示,常牺牲主体对齐精度。本文提出Diptych Prompting,一种新型零样本方法,将主体生成重新理解为利用大规模文本-图像模型中涌现的对偶生成特性进行精确主体对齐的修复任务。该方法将参考图像置于左面板,对右面板进行文本条件修复。通过移除参考图像背景并增强两面板间注意力权重,有效防止内容泄露并提升细节质量。实验表明,本方法显著优于现有零样本图像提示方法,在用户偏好测试中表现更优。此外,该方法还支持主体驱动生成、风格化生成与主体驱动编辑,展现出在多样化图像生成应用中的通用性。
原文摘要 · Abstract (English)
Subject-driven text-to-image generation aims to produce images of a new subject within a desired context by accurately capturing both the visual characteristics of the subject and the semantic content of a text prompt. Traditional methods rely on time- and resource-intensive fine-tuning for subject alignment, while recent zero-shot approaches leverage on-the-fly image prompting, often sacrificing subject alignment. In this paper, we introduce Diptych Prompting, a novel zero-shot approach that reinterprets as an inpainting task with precise subject alignment by leveraging the emergent property of diptych generation in large-scale text-to-image models. Diptych Prompting arranges an incomplete diptych with the reference image in the left panel, and performs text-conditioned inpainting on the right panel. We further prevent unwanted content leakage by removing the background in the reference image and improve fine-grained details in the generated subject by enhancing attention weights between the panels during inpainting. Experimental results confirm that our approach significantly outperforms zero-shot image prompting methods, resulting in images that are visually preferred by users. Additionally, our method supports not only subject-driven generation but also stylized image generation and subject-driven image editing, demonstrating versatility across diverse image generation applications. Project page: https://diptychprompting.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。