不同GPT模型评估视觉描述时各有偏见,评估能力不随通用能力提升而增强。
Understanding AI Evaluation Patterns: How Different GPT Models Assess Vision-Language Descriptions
- 对比三款GPT模型对视觉语言描述的评估策略,发现各具独特评估性格。
- 所有GPT模型均呈现2:1偏向负面评价,但保守性与波动性因模型而异。
- 评估风格是模型固有特性,跨架构对比显示GPT与Gemini策略差异显著。
随着AI系统越来越多地评估其他AI输出,理解其评估行为对防止级联偏差至关重要。本研究分析了NVIDIA的Describe Anything Model生成的视觉语言描述,由GPT-4o、GPT-4o-mini和GPT-5三种变体进行评估,揭示了各模型独特的‘评估性格’及其内在评估策略与偏见。GPT-4o-mini表现出系统性一致性且方差极小,GPT-4o在错误检测上表现优异,而GPT-5则展现出极端保守性且结果波动极大。通过使用Gemini 2.5 Pro作为独立问题生成器的受控实验,验证了这些性格特征为模型内在属性而非人为产物。跨家族分析显示,生成问题的语义相似度表明:GPT模型间聚类紧密、相似度高,而Gemini则表现出显著不同的评估策略。所有GPT模型均一致表现出2:1的偏向负面评估的倾向,但该模式属于模型家族特有,并非所有AI架构共有的普遍现象。研究提示,评估能力并不随通用能力提升而增长,稳健的AI评估需依赖多样化架构视角。
原文摘要 · Abstract (English)
As AI systems increasingly evaluate other AI outputs, understanding their assessment behavior becomes crucial for preventing cascading biases. This study analyzes vision-language descriptions generated by NVIDIA's Describe Anything Model and evaluated by three GPT variants (GPT-4o, GPT-4o-mini, GPT-5) to uncover distinct "evaluation personalities" the underlying assessment strategies and biases each model demonstrates. GPT-4o-mini exhibits systematic consistency with minimal variance, GPT-4o excels at error detection, while GPT-5 shows extreme conservatism with high variability. Controlled experiments using Gemini 2.5 Pro as an independent question generator validate that these personalities are inherent model properties rather than artifacts. Cross-family analysis through semantic similarity of generated questions reveals significant divergence: GPT models cluster together with high similarity while Gemini exhibits markedly different evaluation strategies. All GPT models demonstrate a consistent 2:1 bias favoring negative assessment over positive confirmation, though this pattern appears family-specific rather than universal across AI architectures. These findings suggest that evaluation competence does not scale with general capability and that robust AI assessment requires diverse architectural perspectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。