首个评估人类与大模型文本生成图像提示能力的基准
AtelierEval: Agentic Evaluation of Humans & LLMs as Text-to-Image Prompters

- 构建统一评测框架,覆盖360个专家设计任务
- 提出AtelierJudge评估器,相关性达0.79接近真人水平
- 发现模仿优于规划,适合提示工程与多模态研究者
文本到图像(T2I)系统越来越依赖上游提示生成者(人类或多模态大语言模型,MLLM)将用户意图转化为详细提示。然而现有基准固定提示仅评估T2I模型,未衡量上游提示生成能力。我们提出AtelierEval,首个统一基准,量化360个专家设计任务中的提示能力。基于认知视角,涵盖三类任务,使用真实世界挑战分类法构建任务,并提供面向人类和MLLM的双接口。为实现可扩展、可靠的评估,提出AtelierJudge——一种基于技能、带记忆增强的智能体评估器,对提示-图像对生成主观与客观评分,与人类专家的斯皮尔曼相关性达0.79,接近人类表现。大量实验在4种T2I后端上对比8个MLLM与48名人类用户,验证AtelierEval作为稳健诊断工具的有效性,揭示模仿优于规划,倡导未来提示器向图像增强方向发展。相关工作已开源以支持后续研究。
原文摘要 · Abstract (English)
Text-to-image (T2I) systems increasingly rely on upstream prompters, either humans or multimodal large language models (MLLMs), to translate user intent into detailed prompts. Yet current benchmarks fix the prompt and only evaluate T2I models, leaving the prompting proficiency of this upstream component entirely unmeasured. We introduce AtelierEval, the first unified benchmark that quantifies prompting proficiency across 360 expert-crafted tasks. Grounded in a cognitive view, it spans three task categories and instantiates tasks using a taxonomy of real-world challenges, with a dual interface for both humans and MLLMs. To enable scalable and reliable evaluation, we propose AtelierJudge, a skill-based, memory-augmented agentic evaluator. It produces subjective and objective scores for prompt-image pairs, achieving a Spearman correlation of 0.79 with human experts, approaching human performance. Extensive experiments benchmark 8 MLLMs against 48 human users across 4 T2I backends, validate AtelierEval as a robust diagnostic tool, and reveal the superiority of mimicry over planning, advocating for an image-augmented direction for future prompters. Our work is released to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。