arXiv:2505.11178cs.CVcs.AI2025-05被引 6

用新基准和细粒度反馈提升图文生成对复杂空间关系的准确率

CompAlign: Improving Compositional Text-to-Image Generation with a Complex Benchmark and Fine-Grained Feedback

  • 构建含3个以上主体与复杂3D空间关系的900个测试题
  • 通过分解提示词获取逐项二值反馈,精准量化生成匹配度
  • 基于反馈优化扩散模型,显著提升复杂场景生成能力

当前顶尖文本到图像(T2I)模型虽能生成高分辨率图像,但在表达多对象、属性及空间关系的组合场景时仍存在困难。本文提出CompAlign,一个聚焦3D空间关系评估的挑战性基准,包含900个复杂多主体生成任务,融合数量与3D空间关系,并涵盖多样属性绑定。该基准极具挑战性,涉及3个以上生成主体与复杂3D空间配置。同时提出CompQuest,一种可解释且准确的评估框架,将复杂提示分解为原子子问题,利用多模态大模型(MLLM)提供生成元素各方面的细粒度二值反馈,实现生成图像与组合提示间对齐的精确量化。此外,提出一种对齐框架,以CompQuest反馈作为偏好信号,改进扩散模型的组合图像生成能力。通过可调每图偏好,方法具备良好可扩展性与灵活性。对9个T2I模型的评估显示:(1) 模型在更复杂的3D空间配置下表现显著下降;(2) 开源模型与闭源商业模型之间存在明显性能差距。进一步实证研究表明,使用CompAlign进行模型对齐后,对齐后的扩散模型在组合准确性上取得显著提升,尤其在复杂任务中优于此前方法。

原文摘要 · Abstract (English)

State-of-the-art T2I models are capable of generating high-resolution images given textual prompts. However, they still struggle with accurately depicting compositional scenes that specify multiple objects, attributes, and spatial relations. We present CompAlign, a challenging benchmark with an emphasis on assessing the depiction of 3D-spatial relationships, for evaluating and improving models on compositional image generation. CompAlign consists of 900 complex multi-subject image generation prompts that combine numerical and 3D-spatial relationships with varied attribute bindings. Our benchmark is remarkably challenging, incorporating generation tasks with 3+ generation subjects with complex 3D-spatial relationships. Additionally, we propose CompQuest, an interpretable and accurate evaluation framework that decomposes complex prompts into atomic sub-questions, then utilizes a MLLM to provide fine-grained binary feedback on the correctness of each aspect of generation elements in model-generated images. This enables precise quantification of alignment between generated images and compositional prompts. Furthermore, we propose an alignment framework that uses CompQuest's feedback as preference signals to improve diffusion models' compositional image generation abilities. Using adjustable per-image preferences, our method is easily scalable and flexible for different tasks. Evaluation of 9 T2I models reveals that: (1) models remarkable struggle more with compositional tasks with more complex 3D-spatial configurations, and (2) a noticeable performance gap exists between open-source accessible models and closed-source commercial models. Further empirical study on using CompAlign for model alignment yield promising results: post-alignment diffusion models achieve remarkable improvements in compositional accuracy, especially on complex generation tasks, outperforming previous approaches.

图文生成空间关系评估基准扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。