让多实例图像生成更懂物体间关系与属性
RaDL: Relation-aware Disentangled Learning for Multi-Instance Text-to-Image Generation
- 通过关系注意力捕捉提示中动词表达的物体关系
- 在多个基准上提升位置准确率与属性保留度
- 适合需要精准控制多对象布局的图像生成任务
随着文本到图像(T2I)模型的发展,单个提示生成多个实例成为关键挑战。现有方法虽能生成个体实例位置,但常忽略关系差异和多属性泄露问题。本文提出关系感知解耦学习(RaDL)框架,通过可学习参数增强实例专属属性,并利用全局提示中提取的动作动词生成关系感知图像特征。在COCO-Position、COCO-MIG和DrawBench等基准上的大量实验表明,RaDL显著优于现有方法,在位置准确性、多属性考虑及实例间关系建模方面均有提升。结果证明RaDL是兼顾实例关系与多属性的多实例图像生成有效方案。
原文摘要 · Abstract (English)
With recent advancements in text-to-image (T2I) models, effectively generating multiple instances within a single image prompt has become a crucial challenge. Existing methods, while successful in generating positions of individual instances, often struggle to account for relationship discrepancy and multiple attributes leakage. To address these limitations, this paper proposes the relation-aware disentangled learning (RaDL) framework. RaDL enhances instance-specific attributes through learnable parameters and generates relation-aware image features via Relation Attention, utilizing action verbs extracted from the global prompt. Through extensive evaluations on benchmarks such as COCO-Position, COCO-MIG, and DrawBench, we demonstrate that RaDL outperforms existing methods, showing significant improvements in positional accuracy, multiple attributes consideration, and the relationships between instances. Our results present RaDL as the solution for generating images that consider both the relationships and multiple attributes of each instance within the multi-instance image.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。