让生成图像中的物体保持身份并准确表达关系
DreamRelation: Bridging Customization and Relation Generation
- 分离身份与关系学习,用关键点损失调整物体姿态
- 引入局部特征避免重叠物体混淆,提升关系准确性
- 适用于需要精准物体关系的个性化图像生成场景
定制化图像生成对基于用户提示创建个性化内容至关重要,使大规模文本到图像扩散模型更有效满足个体需求。然而,现有模型常忽略生成图像中定制物体之间的关系。本文提出 DreamRelation 框架,专注于关系感知的定制化图像生成,旨在保留图像提示中的物体身份,同时遵循文本提示指定的关系。通过精心构建的数据集,包含特定关系图像、含身份信息的独立物体图像及文本提示,实现身份与关系解耦学习。提出两个核心模块:一是关键点匹配损失,有效引导模型在需显著姿态调整时生成准确自然的关系;二是利用图像提示的局部特征,更好区分重叠情况下的物体,防止混淆。在自建基准上的大量实验表明,DreamRelation 在多样化物体和关系下,能精确生成关系并保持物体身份。
原文摘要 · Abstract (English)
Customized image generation is essential for creating personalized content based on user prompts, allowing large-scale text-to-image diffusion models to more effectively meet individual needs. However, existing models often neglect the relationships between customized objects in generated images. In contrast, this work addresses this gap by focusing on relation-aware customized image generation, which seeks to preserve the identities from image prompts while maintaining the relationship specified in text prompts. Specifically, we introduce DreamRelation, a framework that disentangles identity and relation learning using a carefully curated dataset. Our training data consists of relation-specific images, independent object images containing identity information, and text prompts to guide relation generation. Then, we propose two key modules to tackle the two main challenges: generating accurate and natural relationships, especially when significant pose adjustments are required, and avoiding object confusion in cases of overlap. First, we introduce a keypoint matching loss that effectively guides the model in adjusting object poses closely tied to their relationships. Second, we incorporate local features of the image prompts to better distinguish between objects, preventing confusion in overlapping cases. Extensive results on our proposed benchmarks demonstrate the superiority of DreamRelation in generating precise relations while preserving object identities across a diverse set of objects and relationships.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。