构建系统化评估框架,提升文生图模型评价的实用性和可解释性
MEF: A Systematic Evaluation Framework for Text-to-Image Models
- 提出涵盖用户场景与文本表达的结构化分类体系
- 结合ELO与MOS实现整体排名与细粒度评分双输出
- 量化分析各维度对用户满意度的贡献,支持模型优化
文生图(T2I)生成技术快速发展,对评估方法提出了更高要求。现有基准多聚焦客观能力维度,缺乏应用场景视角,外部有效性受限;且评估常依赖ELO整体排名或MOS分维打分,两者均存在固有缺陷且解释性不足。为此,本文提出魔法评估框架(MEF),构建包含用户场景、元素、组合方式与文本形式的结构化分类体系,形成可支持标签级评估的Magic-Bench-377数据集,确保场景与能力覆盖均衡。在此基础上,融合ELO与维度MOS,分别实现模型整体排名与细粒度评估,并通过多元逻辑回归定量分析各维度对用户满意度的贡献。基于MEF对主流T2I模型的评估,获得领先模型排行榜及其关键特征。框架与Magic-Bench-377已开源,推动视觉生成模型评估研究发展。
原文摘要 · Abstract (English)
Rapid advances in text-to-image (T2I) generation have raised higher requirements for evaluation methodologies. Existing benchmarks center on objective capabilities and dimensions, but lack an application-scenario perspective, limiting external validity. Moreover, current evaluations typically rely on either ELO for overall ranking or MOS for dimension-specific scoring, yet both methods have inherent shortcomings and limited interpretability. Therefore, we introduce the Magic Evaluation Framework (MEF), a systematic and practical approach for evaluating T2I models. First, we propose a structured taxonomy encompassing user scenarios, elements, element compositions, and text expression forms to construct the Magic-Bench-377, which supports label-level assessment and ensures a balanced coverage of both user scenarios and capabilities. On this basis, we combine ELO and dimension-specific MOS to generate model rankings and fine-grained assessments respectively. This joint evaluation method further enables us to quantitatively analyze the contribution of each dimension to user satisfaction using multivariate logistic regression. By applying MEF to current T2I models, we obtain a leaderboard and key characteristics of the leading models. We release our evaluation framework and make Magic-Bench-377 fully open-source to advance research in the evaluation of visual generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。