分析现有自动评分指标,发现单一指标难拟人,建议用混合评分提升评估质量。
A Step Towards Mixture of Grader: Statistical Analysis of Existing Automatic Evaluation Metrics
- 通过相关性分析发现不同指标间高度相似
- 无单一指标能全面贴近人工评分
- 提出混合评分器可改善自动评估效果
开源模型和问答数据集的爆发式增长凸显了自动化问答评估的重要性。本文研究了现有评估指标的统计特性,以更深入理解其局限性。通过测量各评估指标与类人评分之间的相关系数,发现:(1) 在问题类型(如单字、单短语等)层面,现有指标之间具有高度相关性;(2) 没有单一指标能够充分逼近人类评分。作为潜在解决方案,本文探讨了构建混合评分器(Mixture of Grader)在提升自动问答评估质量方面的可能性。
原文摘要 · Abstract (English)
The explosion of open-sourced models and Question-Answering (QA) datasets emphasizes the importance of automated QA evaluation. We studied the statistics of the existing evaluation metrics for a better understanding of their limitations. By measuring the correlation coefficients of each evaluation metric concerning human-like evaluation score, we observed the following: (1) existing metrics have a high correlation among them concerning the question type (e.g., single word, single phrase, etc.), (2) no single metric can adequately estimate the human-like evaluation. As a potential solution, we discuss how a Mixture Of Grader could potentially improve the auto QA evaluator quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。