构建带结构标注的大规模图文数据集,提升复杂场景生成质量
LAION-SG: An Enhanced Large-Scale Dataset for Training Complex Image-Text Models with Structural Annotations
- 基于场景图结构标注,精确描述多对象属性与关系
- 新模型在复杂场景生成上显著优于现有模型
- 提供数据、模型和评测基准,推动领域发展
文本到图像生成虽已取得显著进展,但在涉及多个对象及复杂关系的组合生成任务中表现下降。问题根源在于现有图文数据集缺乏精确的跨对象关系标注。为此,我们构建了LAION-SG——一个大规模高质量场景图(SG)结构标注数据集,精准描述多对象的属性与相互关系,有效表征复杂场景语义结构。基于该数据集,我们训练了新基础模型SDXL-SG,将结构信息融入生成过程。大量实验证明,基于LAION-SG训练的模型在复杂场景生成任务上显著优于现有模型。同时,我们提出CompSG-Bench评测基准,用于评估组合图像生成能力,建立该领域的新的评价标准。相关标注数据、处理代码、基础模型及评测协议均开源:https://github.com/mengcye/LAION-SG。
原文摘要 · Abstract (English)
Recent advances in text-to-image (T2I) generation have shown remarkable success in producing high-quality images from text. However, existing T2I models show decayed performance in compositional image generation involving multiple objects and intricate relationships. We attribute this problem to limitations in existing datasets of image-text pairs, which lack precise inter-object relationship annotations with prompts only. To address this problem, we construct LAION-SG, a large-scale dataset with high-quality structural annotations of scene graphs (SG), which precisely describe attributes and relationships of multiple objects, effectively representing the semantic structure in complex scenes. Based on LAION-SG, we train a new foundation model SDXL-SG to incorporate structural annotation information into the generation process. Extensive experiments show advanced models trained on our LAION-SG boast significant performance improvements in complex scene generation over models on existing datasets. We also introduce CompSG-Bench, a benchmark that evaluates models on compositional image generation, establishing a new standard for this domain. Our annotations with the associated processing code, the foundation model and the benchmark protocol are publicly available at https://github.com/mengcye/LAION-SG.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。