arXiv:2501.15595cs.CV2025-01EMNLP被引 29

用自适应评分标准提升大模型评估精度,让机器评卷更像真人。

SedarEval: Automated Evaluation using Self-Adaptive Rubrics

  • 为每道题定制动态评分标准,模仿人类评分逻辑。
  • 在1000道跨领域题目上,评分一致性超越GPT-4。
  • 适合需要高精度自动评估的研究者与开发者使用。

LLM作为评分员的评估范式因显著降低人力与时间成本而广受欢迎。该方法利用一个或多个大语言模型(LLMs)评估其他LLMs输出质量。然而,现有方法依赖通用评分标准,未能考虑每道题及其解题过程的特殊性,影响评估的精确性与稳定性。受人类考试评分流程启发,我们提出基于自适应评分标准的新评估范式。具体而言,为每道题构建详尽的评分规则,以结构化方式呈现主要与次要评分项及扣分点,模拟人类评分员的分析过程。在此基础上,我们进一步开发了新基准SedarEval,涵盖长尾知识、数学、编程与逻辑推理等多个领域,包含1000道精心设计的问题,每题均有独立的自适应评分标准。为实现评估自动化,我们训练了一个专用评分语言模型(evaluator LM),取代人工评分员。使用相同训练数据,该模型在与人工评分结果的一致性上优于其他范式,包括GPT-4,凸显本方法的优越性与高效性。数据集已公开于https://github.com/wwn1233/sedareval。

原文摘要 · Abstract (English)

The evaluation paradigm of LLM-as-judge gains popularity due to its significant reduction in human labor and time costs. This approach utilizes one or more large language models (LLMs) to assess the quality of outputs from other LLMs. However, existing methods rely on generic scoring rubrics that fail to consider the specificities of each question and its problem-solving process, compromising precision and stability in assessments. Inspired by human examination scoring processes, we propose a new evaluation paradigm based on self-adaptive rubrics. Specifically, we create detailed scoring rubrics for each question, capturing the primary and secondary criteria in a structured format of scoring and deduction points that mimic a human evaluator's analytical process. Building on this paradigm, we further develop a novel benchmark called SedarEval, which covers a range of domains including long-tail knowledge, mathematics, coding, and logical reasoning. SedarEval consists of 1,000 meticulously crafted questions, each with its own self-adaptive rubric. To further streamline the evaluation, we train a specialized evaluator language model (evaluator LM) to supplant human graders. Using the same training data, our evaluator LM achieves a higher concordance rate with human grading results than other paradigms, including GPT-4, highlighting the superiority and efficiency of our approach. We release our dataset at https://github.com/wwn1233/sedareval.

自动评估评分标准大模型评测LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。