arXiv:2505.24787cs.CVcs.CL2025-05被引 32

构建复杂指令图像生成的评测基准与智能代理框架

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation

  • 设计9维500条复杂提示的评测基准LongBench-T2I
  • 提出Plan2Gen框架,无需训练即可解析复杂指令
  • 开发多维度自动化评估工具,超越传统度量

文本到图像(T2I)生成技术虽已取得进展,但面对包含多个对象、属性和空间关系的复杂指令时仍表现不佳。现有评测基准主要关注文本-图像对齐,难以衡量复杂提示的执行能力。为此,我们提出LongBench-T2I,一个包含500条精心设计提示的综合性评测基准,覆盖九种视觉评估维度,全面评估模型对复杂指令的理解与生成能力。此外,我们提出Plan2Gen代理框架,利用大语言模型解析并分解复杂提示,引导现有T2I模型生成更符合要求的图像,且无需额外训练。针对现有评估指标(如CLIPScore)在复杂指令评估中的不足,我们引入一套多维度自动化评估工具,提升评估精度。数据与代码已公开于https://github.com/yczhou001/LongBench-T2I。

原文摘要 · Abstract (English)

Recent advancements in text-to-image (T2I) generation have enabled models to produce high-quality images from textual descriptions. However, these models often struggle with complex instructions involving multiple objects, attributes, and spatial relationships. Existing benchmarks for evaluating T2I models primarily focus on general text-image alignment and fail to capture the nuanced requirements of complex, multi-faceted prompts. Given this gap, we introduce LongBench-T2I, a comprehensive benchmark specifically designed to evaluate T2I models under complex instructions. LongBench-T2I consists of 500 intricately designed prompts spanning nine diverse visual evaluation dimensions, enabling a thorough assessment of a model's ability to follow complex instructions. Beyond benchmarking, we propose an agent framework (Plan2Gen) that facilitates complex instruction-driven image generation without requiring additional model training. This framework integrates seamlessly with existing T2I models, using large language models to interpret and decompose complex prompts, thereby guiding the generation process more effectively. As existing evaluation metrics, such as CLIPScore, fail to adequately capture the nuances of complex instructions, we introduce an evaluation toolkit that automates the quality assessment of generated images using a set of multi-dimensional metrics. The data and code are released at https://github.com/yczhou001/LongBench-T2I.

图像生成复杂指令评测基准智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。