arXiv:2502.18339cs.CLcs.LG2025-02被引 6

发现传统评测能准确预测人类对聊天模型的偏好,节省大量人工评估成本。

Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks

  • 用160个标准NLP任务对比4个聊天模型的表现与真人评价
  • 多数任务与人类偏好强相关,可预测用户满意度
  • 适用于想降低人工评测成本的研究者和产品团队

对话式语言模型的爆发促使评估方式从传统的自然语言处理(NLP)基准转向昂贵、耗时且噪声较大的人工评价,但两者之间的关系尚不清晰。本文对四个Chat Llama 2模型开展大规模研究,比较其在160个标准NLP基准(如MMLU、ARC、BIG-Bench Hard)上的表现,与超过2000名标注者提供的11000余条单轮及2000余条多轮对话中的人类偏好。结果显示:大多数NLP基准与人类评价高度相关,表明自动化指标可作为人类偏好的可靠预测工具;而对抗性欺骗、安全性等三项人类评价与基准呈负相关,另有两项无显著相关性。通过过参数化线性回归,我们进一步证明:仅凭NLP得分即可准确预测不同规模模型的人类评价,为减少高成本人工标注提供了可行路径。总体而言,研究证实了传统基准的持续价值,并揭示如何利用它们预判真实用户满意度,为当前对话式AI的评估需求提供新思路。

原文摘要 · Abstract (English)

The explosion of high-performing conversational language models (LMs) has spurred a shift from classic natural language processing (NLP) benchmarks to expensive, time-consuming and noisy human evaluations - yet the relationship between these two evaluation strategies remains hazy. In this paper, we conduct a large-scale study of four Chat Llama 2 models, comparing their performance on 160 standard NLP benchmarks (e.g., MMLU, ARC, BIG-Bench Hard) against extensive human preferences on more than 11k single-turn and 2k multi-turn dialogues from over 2k human annotators. Our findings are striking: most NLP benchmarks strongly correlate with human evaluations, suggesting that cheaper, automated metrics can serve as surprisingly reliable predictors of human preferences. Three human evaluations, such as adversarial dishonesty and safety, are anticorrelated with NLP benchmarks, while two are uncorrelated. Moreover, through overparameterized linear regressions, we show that NLP scores can accurately predict human evaluations across different model scales, offering a path to reduce costly human annotation without sacrificing rigor. Overall, our results affirm the continued value of classic benchmarks and illuminate how to harness them to anticipate real-world user satisfaction - pointing to how NLP benchmarks can be leveraged to meet evaluation needs of our new era of conversational AI.

语言模型人类评估自动评测对话AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。