用3D数据生成多图合成数据,提升文本到图像定制效果
Generating Multi-Image Synthetic Data for Text-to-Image Customization
- 用3D数据生成同一物体在不同光照背景下的多图合成数据
- 训练的编码器模型在标准评测上超越现有方法,提升生成质量
- 适合需要高质量定制化图像生成的研究者与开发者
文本到图像模型的定制化允许用户插入新概念或对象,并在未见场景中生成。现有方法或依赖昂贵的测试时优化,或在单图数据集上训练编码器而缺乏多图监督,限制了图像质量。本文提出一种简单方法:首先利用现有文本到图像模型和3D数据集,构建包含同一物体在不同光照、背景和姿态下的多图合成定制数据集(SynCD);接着基于该数据集训练一个编码器模型,通过共享注意力机制融合参考图像中的细粒度视觉细节;最后提出一种推理技术,对文本与图像引导向量进行归一化,缓解生成图像过曝问题。大量实验表明,基于SynCD训练的编码器模型结合所提推理算法,在标准定制化基准测试上优于现有方法。
原文摘要 · Abstract (English)
Customization of text-to-image models enables users to insert new concepts or objects and generate them in unseen settings. Existing methods either rely on comparatively expensive test-time optimization or train encoders on single-image datasets without multi-image supervision, which can limit image quality. We propose a simple approach to address these challenges. We first leverage existing text-to-image models and 3D datasets to create a high-quality Synthetic Customization Dataset (SynCD) consisting of multiple images of the same object in different lighting, backgrounds, and poses. Using this dataset, we train an encoder-based model that incorporates fine-grained visual details from reference images via a shared attention mechanism. Finally, we propose an inference technique that normalizes text and image guidance vectors to mitigate overexposure issues in sampled images. Through extensive experiments, we show that our encoder-based model, trained on SynCD, and with the proposed inference algorithm, improves upon existing encoder-based methods on standard customization benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。