为文生图模型设计了更贴近真实创作的评估体系。
Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation

- 联合艺术家构建分层评估框架,覆盖真实还原与创意生成
- 用56个可验证标准评分,区分顶尖文生图模型性能差异
- 适合追求高质量图像生成的创作者和研发团队使用
文生图技术已从基础图像合成演变为专业创作流程的核心能力,单纯依赖文本对齐已无法满足用户对真实世界还原和真实创意表达的需求。现有评测基准仍停留在基础指标,难以捕捉真实艺术实践中关键的细微能力,导致难以可靠区分顶尖文生图模型。为此,我们提出Qwen-Image-Bench,一个由专业艺术家共同设计、基于真实创作场景的以创作者为中心的评测基准。该基准在传统评估基础上引入两个应用驱动维度:真实世界保真度与创意生成能力。基于专业艺术创作中的分阶段推理逻辑,将五项核心支柱组织为自上而下的层级分类体系,进一步分解为23个二级子能力与56个三级可验证标准。为确保覆盖面,我们精心策划了1000个分层提示词,每个提示词同时涉及多个支柱中超过四个细粒度方面。我们训练了统一判别模型Q-Judger,基于Qwen3.6-27B,由来自全球艺术学院的80位专业标注员在盲标与三重评审协议下监督训练,对每张图像在全部56个可验证维度进行评分,生成细粒度、基于标准且可溯源的诊断结果,而非单一模糊得分。实证表明,Qwen-Image-Bench能有效区分领先文生图模型,在真实世界保真度与创意生成这两个现有基准提供极少洞察的维度上实现最大区分度,同时为生产级文生图开发提供可信优化信号。
原文摘要 · Abstract (English)
Text-to-Image generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple text-image alignment can no longer satisfy users' pressing demands for faithful real-world reconstruction and genuine creative expression. Existing benchmarks, however, remain anchored in these foundational criteria and do not yet capture the nuanced capabilities that matter in authentic artistic practice, making it difficult to reliably distinguish state-of-the-art T2I models. To address the gap, we introduce Qwen-Image-Bench, a creator-centric benchmark co-designed with professional artists and grounded in real-world creation scenarios. Qwen-Image-Bench enriches conventional evaluation with two application-driven dimensions: Real-world Fidelity and Creative Generation. Drawing on the staged reasoning inherent in professional artistic workflows, we organize these five pillars into a top-down hierarchical taxonomy that further decomposes into 23 second-level sub-capabilities and 56 third-level verifiable rubrics. To ensure broad coverage, we curate 1000 stratified prompts with each prompt jointly exercising more than four fine-grained facets across multiple pillars. We train a unified judge model Q-Judger based on Qwen3.6-27B, supervised by 80 professional annotators from global art academies under blind labeling and triple-review protocols, that scores every image across all 56 verifiable facets, producing fine-grained, rubric-grounded, and fully attributable diagnostics rather than a single opaque score. Empirically, Qwen-Image-Bench reliably distinguishes leading T2I models, achieving the greatest separation on the two application-driven dimensions of Real-world Fidelity and Creative Generation where existing benchmarks provide little insight, while also providing a trustworthy optimization signal for production-level T2I development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。