发现图像评估指标偏爱典型图像,而非准确匹配提示的图像。
Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics
- 设计对照实验,用语义正确与典型但错误图像对比测试评估指标。
- 主流评估方法在70%以上案例中偏好错误但典型的图像。
- 适合关注生成模型真实性的研究者和评测系统开发者。
自动评估指标广泛用于文本到图像模型的评测,常取代人工判断进行基准测试、模型选择和大规模数据过滤。然而,这些指标可能奖励外观或社会上典型的图像,而非忠实满足提示的图像。我们识别出一种系统性盲点——原型性偏差:评估指标可能更偏好语义错误但视觉或社会上典型的图像,而非语义正确但不够典型的图像。为此,我们构建了PROTOBIAS,一个跨动物、物体和人口群体的受控诊断基准,其中语义正确的图像与包含单一可控语义偏差的典型对抗图像进行对比。该基准基于原型理论和社会类别原型性,采用多种提示生成器、图像生成器和独立的视觉语言模型(VLM)过滤器,并通过提示质量、人工标注和图像质量控制进行验证。使用PROTOBIAS,我们发现嵌入、奖励、基于VQA及VLM作为裁判的常用指标在多数情况下无法区分正确与典型错误图像,而人类判断仍更忠实于语义正确性。我们进一步提出PROTOSCORE,一种轻量级对比训练的评估器,作为初步缓解基线。PROTOBIAS为测量原型性驱动的评估失败提供了聚焦基准,并推动更语义忠实的文本到图像评估器发展。
原文摘要 · Abstract (English)
Automatic metrics are widely used to evaluate text-to-image models, often replacing human judgment in benchmarking, model selection, and large-scale data filtering. Yet they may reward images that look plausible or prototypical rather than images that faithfully satisfy the prompt. We identify prototypicality bias as a systematic blindspot in multimodal evaluation: metrics can prefer a semantically incorrect but visually or socially prototypical image over a correct but less prototypical one. We introduce PROTOBIAS, a controlled diagnostic benchmark across Animals, Objects, and Demography, where semantically correct images are contrasted with plausible prototypical adversaries containing a single controlled semantic violation. Grounded in prototype theory and social-category prototypicality, PROTOBIAS is constructed with multiple prompt generators, image generators, and independent VLM filters, and validated through prompt-quality, human-annotation, and image-quality controls. Using PROTOBIAS, we show that widely used embedding, reward, VQA-based, and VLM-as-judge metrics frequently fail these contrasts, while human judgments remain more faithful to semantic correctness. We further introduce PROTOSCORE, a lightweight contrastively trained evaluator, as an initial mitigation baseline. PROTOBIAS provides a focused benchmark for measuring prototypicality-driven metric failures and developing more semantically faithful T2I evaluators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。