arXiv:2505.16915cs.CVcs.AI2025-05中稿 · ICML被引 7

评测文生图模型对长提示的处理能力,发现现有模型在细节和结构上表现不佳。

DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?

  • 构建包含284.89词长提示的基准测试集,涵盖角色属性、位置、场景与交互关系
  • 实测显示扩散模型在细节密集时出现属性泄露,编码器难保持语法依赖
  • 证明高保真生成需同时扩大提示长度限制与进行长提示训练

尽管近期文生图(T2I)模型在简短描述下表现优异,但在专业应用所需的长而复杂的提示下仍表现不佳。本文提出DetailMaster,一个全面评估T2I模型在长提示下能力的基准,包含专家验证的平均284.89词长提示,以及自动化数据构建流程与评估工作流。该基准涵盖四大关键维度:角色属性、结构化角色位置、多维场景属性与空间/交互关系。对多种通用及长提示优化模型的评测揭示了显著性能瓶颈:弱编码器难以维持提示内的语法依赖,扩散模型在细节密集条件下易出现属性泄露。通过在不同约束下的受控消融实验,进一步表明高保真生成需要提示长度扩展与长提示训练的协同作用。我们开源数据集与代码,以推动长提示驱动的文生图技术发展。

原文摘要 · Abstract (English)

While recent Text-to-Image (T2I) models show impressive capabilities in synthesizing images from brief descriptions, they struggle with the long, detailed prompts required for professional applications. We present DetailMaster, a comprehensive benchmark for evaluating T2I capabilities on long prompts with complex compositional requirements, accompanied by an automated data construction pipeline and an evaluation workflow. Comprising expert-validated prompts averaging 284.89 tokens, our benchmark introduces four critical evaluation dimensions: Character Attributes, Structured Character Locations, Multi-Dimensional Scene Attributes, and Spatial/Interactive Relationships. Evaluations on various general-purpose and long-prompt-optimized models reveal critical performance limitations, showing that weak encoders struggle to preserve syntactic dependencies within prompts and diffusion models suffer from attribute leakage under detail-intensive conditions. Through a controlled ablation study under varying constraints, we further show that high-fidelity generation requires a synergistic combination of expanded prompt limits and long-prompt training. We open-source our dataset and code to foster progress in long-prompt-driven T2I generation.

文生图长提示基准评测扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。