改进布局到图像生成的注意力机制与评估方法
Rethinking The Training And Evaluation of Rich-Context Layout-to-Image Generation
- 设计区域交叉注意力模块增强复杂布局表征
- 提出新指标评估开放词汇下的生成性能
- 用户研究验证指标与人类偏好一致
生成模型在图像生成领域取得显著进展,广泛应用于图像编辑、补全和视频编辑。布局到图像(L2I)生成是其中一类特殊任务,通过预定义的对象布局引导生成过程。本文提出一种新型区域交叉注意力模块,显著提升布局区域的表征能力,尤其在处理高度复杂且细节丰富的文本描述时表现更优。此外,尽管当前开放词汇L2I方法在开放集环境下训练,其评估却常在封闭集环境中进行。为弥合这一差距,本文提出两项适用于开放词汇场景的评估指标,并通过全面的用户研究验证了这些指标与人类偏好的一致性。
原文摘要 · Abstract (English)
Recent advancements in generative models have significantly enhanced their capacity for image generation, enabling a wide range of applications such as image editing, completion and video editing. A specialized area within generative modeling is layout-to-image (L2I) generation, where predefined layouts of objects guide the generative process. In this study, we introduce a novel regional cross-attention module tailored to enrich layout-to-image generation. This module notably improves the representation of layout regions, particularly in scenarios where existing methods struggle with highly complex and detailed textual descriptions. Moreover, while current open-vocabulary L2I methods are trained in an open-set setting, their evaluations often occur in closed-set environments. To bridge this gap, we propose two metrics to assess L2I performance in open-vocabulary scenarios. Additionally, we conduct a comprehensive user study to validate the consistency of these metrics with human preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。