从因果视角设计无偏评估,揭示大模型评测中的隐藏偏差
Unbiased Evaluation of Large Language Models from a Causal Perspective
- 用因果分析框架识别评测中代理生成问题的两种偏差
- 新评估协议在多个数据集上暴露当前大模型显著不足
- 适合关注评测公平性与模型真实能力的研究者
大语言模型评测中的基准污染问题日益严重。以往的‘代理作为评测者’方法虽通过代理生成问题缓解此问题,但其自身仍存在未被充分研究的偏差。本文提出一种评估偏差的理论框架,通过在最小化代理评测设置下设计探测任务,识别出两类关键偏差。为此,我们提出无偏评测器(Unbiased Evaluator),可提供更全面、无偏且可解释的模型评估结果。大量实验表明,当前大模型仍有巨大提升空间。此外,该方法不仅有力证明了基准污染的存在,还能给出可解释的评估结论。
原文摘要 · Abstract (English)
Benchmark contamination has become a significant concern in the LLM evaluation community. Previous Agents-as-an-Evaluator address this issue by involving agents in the generation of questions. Despite their success, the biases in Agents-as-an-Evaluator methods remain largely unexplored. In this paper, we present a theoretical formulation of evaluation bias, providing valuable insights into designing unbiased evaluation protocols. Furthermore, we identify two type of bias in Agents-as-an-Evaluator through carefully designed probing tasks on a minimal Agents-as-an-Evaluator setup. To address these issues, we propose the Unbiased Evaluator, an evaluation protocol that delivers a more comprehensive, unbiased, and interpretable assessment of LLMs.Extensive experiments reveal significant room for improvement in current LLMs. Additionally, we demonstrate that the Unbiased Evaluator not only offers strong evidence of benchmark contamination but also provides interpretable evaluation results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。