GLIDER用可解释评分提升大模型评估的精度与透明度。
GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking
- 基于30亿参数设计,支持任意自定义标准评分
- 在FLASK上相关性超越GPT-4o,性能接近17倍大的模型
- 支持细粒度打分、多语言推理与标注高亮,适合研究者使用
大模型作为评判者(LLM-as-judge)正被广泛用于自动化评估模型输出。尽管在特定任务中表现良好,封闭源代码的大模型在真实场景应用中因缺乏细粒度指标和可解释性而存在明显缺陷,而专用评估模型又缺乏跨领域泛化能力。我们提出GLIDER,一个30亿参数的评估型大模型,可对任意文本输入及其上下文按用户自定义标准进行评分。GLIDER在FLASK数据集上的皮尔逊相关性高于GPT-4o,且性能达到比自身大17倍模型的水平。它支持细粒度评分、多语言推理、片段高亮,并在685个领域和183项标准上训练。定性分析显示,其评分与人类判断高度一致,人类同意率达91.3%。我们已开源GLIDER以推动后续研究。
原文摘要 · Abstract (English)
The LLM-as-judge paradigm is increasingly being adopted for automated evaluation of model outputs. While LLM judges have shown promise on constrained evaluation tasks, closed source LLMs display critical shortcomings when deployed in real world applications due to challenges of fine grained metrics and explainability, while task specific evaluation models lack cross-domain generalization. We introduce GLIDER, a powerful 3B evaluator LLM that can score any text input and associated context on arbitrary user defined criteria. GLIDER shows higher Pearson's correlation than GPT-4o on FLASK and greatly outperforms prior evaluation models, achieving comparable performance to LLMs 17x its size. GLIDER supports fine-grained scoring, multilingual reasoning, span highlighting and was trained on 685 domains and 183 criteria. Extensive qualitative analysis shows that GLIDER scores are highly correlated with human judgments, with 91.3% human agreement. We have open-sourced GLIDER to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。