收集200万张图的人类偏好数据,全面评估AI绘图模型表现
Finding the Subjective Truth: Collecting 2 Million Votes for Comprehensive Gen-AI Model Evaluation
- 通过全球众包平台收集人类对图像的主观评价
- 覆盖4512张图、200万条标注,对比四大主流模型性能
- 多样化的评审者群体降低偏见,结果更具代表性
高效评估文本到图像模型的表现困难,因其本质上依赖主观判断与人类偏好,难以进行模型间比较和当前技术水平量化。借助Rapidata技术,我们提出一种高效的注释框架,从多元化的全球标注者池中获取人类反馈。本研究共收集超过200万条标注,涵盖4,512张图像,评估了DALL-E 3、Flux.1、MidJourney和Stable Diffusion四款主流模型在风格偏好、连贯性及文本-图像对齐方面的表现。结果表明,该方法可实现基于大规模标注者的综合模型排名,且标注者群体结构反映世界人口多样性,显著降低偏差风险。
原文摘要 · Abstract (English)
Efficiently evaluating the performance of text-to-image models is difficult as it inherently requires subjective judgment and human preference, making it hard to compare different models and quantify the state of the art. Leveraging Rapidata's technology, we present an efficient annotation framework that sources human feedback from a diverse, global pool of annotators. Our study collected over 2 million annotations across 4,512 images, evaluating four prominent models (DALL-E 3, Flux.1, MidJourney, and Stable Diffusion) on style preference, coherence, and text-to-image alignment. We demonstrate that our approach makes it feasible to comprehensively rank image generation models based on a vast pool of annotators and show that the diverse annotator demographics reflect the world population, significantly decreasing the risk of biases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。