评估大模型作为评分器的一致性,发现强模型也不一定可靠。
Evaluating the Consistency of LLM Evaluators
- 从自一致性和跨尺度一致性两方面测试大模型评分
- 发现强大专有模型在不同评分标准下仍不一致
- 提醒研究者评估大模型时需关注一致性而非仅看性能
大语言模型(LLMs)在快速、低成本方面展现出作为通用评分器的潜力。尽管其与人工标注的相关性已被广泛研究,但作为评分器的一致性仍缺乏足够关注,引发对其可靠性担忧。本文对开源与专有模型在不同评分尺度和判别粒度下的自一致性(SC)与跨尺度一致性(IC)进行了系统研究。全面分析表明,强大的专有模型未必是稳定的评分者,凸显在评估大模型评分能力时必须考虑一致性问题。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown potential as general evaluators along with the evident benefits of speed and cost. While their correlation against human annotators has been widely studied, consistency as evaluators is still understudied, raising concerns about the reliability of LLM evaluators. In this paper, we conduct extensive studies on the two aspects of consistency in LLM evaluations, Self-Consistency (SC) and Inter-scale Consistency (IC), on different scoring scales and criterion granularity with open-source and proprietary models. Our comprehensive analysis demonstrates that strong proprietary models are not necessarily consistent evaluators, highlighting the importance of considering consistency in assessing the capability of LLM evaluators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。