发现大模型评估时标签会引发严重偏见,影响判断结果。
Quantifying Label-Induced Bias in Large Language Model Self- and Cross-Evaluations
- 用四种标签条件测试三模型互评,控制作者身份
- 错误标签可使评分变化达50个百分点,真实标签下自评偏高或偏低
- 揭示模型身份标签扭曲评估结果,建议采用盲评机制
大型语言模型(LLMs)越来越多地被用于文本质量评估,但其判断的可靠性仍待深入探讨。本研究考察了ChatGPT、Gemini和Claude三款主流模型在自评与互评中存在系统性偏见。设计对照实验,让各模型撰写的博客文章在四种标注条件下(无署名、真实署名、两种虚假署名)由三模型进行评估。评估采用整体偏好投票与细粒度质量评分(连贯性、信息量、简洁性),所有分数标准化为百分比以便比较。结果显示显著的评价不对称:'Claude'标签始终提升得分,而'Gemini'标签则系统性拉低分数;虚假署名常导致偏好排名反转,投票结果变动高达50个百分点,质量评分波动达12个百分点。值得注意的是,Gemini在真实标签下表现出严重自我贬低,而Claude则强化自我偏好。结果表明,模型身份认知会显著扭曲高层级判断与细粒度评估,与内容质量无关。研究质疑了以大模型为评判者的可行性,强调需采用盲评协议和多模型验证框架,保障自动化文本评估与基准测试的公平性与有效性。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed as evaluators of text quality, yet the validity of their judgments remains underexplored. This study investigates systematic bias in self- and cross-model evaluations across three prominent LLMs: ChatGPT, Gemini, and Claude. We designed a controlled experiment in which blog posts authored by each model were evaluated by all three models under four labeling conditions: no attribution, true attribution, and two false-attribution scenarios. Evaluations employed both holistic preference voting and granular quality ratings across three dimensions Coherence, Informativeness, and Conciseness with all scores normalized to percentages for direct comparison. Our findings reveal pronounced asymmetries in model judgments: the "Claude" label consistently elevated scores regardless of actual authorship, while the "Gemini" label systematically depressed them. False attribution frequently reversed preference rankings, producing shifts of up to 50 percentage points in voting outcomes and up to 12 percentage points in quality ratings. Notably, Gemini exhibited severe self-deprecation under true labels, while Claude demonstrated intensified self-preference. These results demonstrate that perceived model identity can substantially distort both high-level judgments and fine-grained quality assessments, independent of content quality. Our findings challenge the reliability of LLM-as-judge paradigms and underscore the critical need for blind evaluation protocols and diverse multi-model validation frameworks to ensure fairness and validity in automated text evaluation and LLM benchmarking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。