构建10万条图文内容评价数据集,提升生成质量与对齐度评估能力。
Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content
- 基于百万级人工标注,构建图文生成质量与对齐评估数据集。
- 提出Q-Eval-Score模型,在视觉质量和长文本对齐上均表现更优。
- 适合图像/视频生成、评估模型研发人员使用。
评估文本到视觉内容的核心在于视觉质量和语义对齐。尽管已有客观评估模型进展,其性能仍依赖于人工标注的规模与质量。根据缩放定律,增加标注数量可系统性提升模型表现。为此,我们提出Q-Eval-100K数据集,包含10万条图文和视频样本的96万条人类标注的平均意见分数(MOS),涵盖6万张图像和4万段视频。利用该数据集与上下文提示,我们设计了统一评估模型Q-Eval-Score,特别优化长文本提示下的对齐评估能力。实验表明,Q-Eval-Score在视觉质量和对齐任务中均优于现有方法,并在多个基准上展现良好泛化性。数据与代码将开源。
原文摘要 · Abstract (English)
Evaluating text-to-vision content hinges on two crucial aspects: visual quality and alignment. While significant progress has been made in developing objective models to assess these dimensions, the performance of such models heavily relies on the scale and quality of human annotations. According to Scaling Law, increasing the number of human-labeled instances follows a predictable pattern that enhances the performance of evaluation models. Therefore, we introduce a comprehensive dataset designed to Evaluate Visual quality and Alignment Level for text-to-vision content (Q-EVAL-100K), featuring the largest collection of human-labeled Mean Opinion Scores (MOS) for the mentioned two aspects. The Q-EVAL-100K dataset encompasses both text-to-image and text-to-video models, with 960K human annotations specifically focused on visual quality and alignment for 100K instances (60K images and 40K videos). Leveraging this dataset with context prompt, we propose Q-Eval-Score, a unified model capable of evaluating both visual quality and alignment with special improvements for handling long-text prompt alignment. Experimental results indicate that the proposed Q-Eval-Score achieves superior performance on both visual quality and alignment, with strong generalization capabilities across other benchmarks. These findings highlight the significant value of the Q-EVAL-100K dataset. Data and codes will be available at https://github.com/zzc-1998/Q-Eval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。