自生成评测让大模型自我偏袒,得分虚高。
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
- 用大模型生成测试题和评分,导致结果偏爱自身。
- 即使控制多样性,模型风格仍使自己得分虚高。
- 适合关注评测可信度的研究者与开发者。
随着大模型迅速填满现有评测集,利用大模型自动创建评测(即‘大模型作为评测集’)成为低成本替代人工标注的热门方法:由模型生成测试输入(大模型作为测试集),并评估输出结果(大模型作为评估器)。我们发现该范式存在根本性问题:生成的评测集系统性地偏袒创建它的模型。以机器翻译为主要测试场景,我们发现自偏倚源于两大因素——大模型作为测试集与作为评估器——其叠加效应显著放大偏差。关键在于,即便测试数据显式控制多样性,各模型隐含的写作风格仍导致输出同质化,进而抬高自身得分。通过引入新多样性度量指标提升源文本多样性,可部分缓解该偏倚。自偏倚强到足以使每个模型都排在自己首位,压倒同行共识排序。该现象在开放生成任务的Chatbot Arena上亦得到验证。
原文摘要 · Abstract (English)
As LLMs rapidly saturate existing benchmarks, automated benchmark creation using LLMs (LLM as a benchmark) where a model generates test inputs (LLM as a testset) and evaluates outputs (LLM as an evaluator) has gained traction as a cheap alternative to human curation. We show that this paradigm has a fundamental problem: LLM-generated benchmarks systematically favor the model that created them. Using machine translation as our primary testbed, we find that self bias arises from two compounding sources, LLM as a testset and LLM as an evaluator, and their combination amplifies the effect. Crucially, even when test data is generated with explicit diversity controls, each modelś implicit stylistic tendencies produce homogeneous, model-specific outputs that inflate its own scores. Increasing source text diversity, using our proposed diversity metric, partially mitigates this bias. Self bias is strong enough to cause each model to rank itself first, overriding the peer consensus ordering. We confirm that the phenomenon extends to open-ended generation on the Chatbot Arena task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。