LLM做评判需人类参考,否则易误判。
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
- 用金融专家标注的160道题测试LLM评判能力
- 无正确答案参考时,LLM评判仅在自己答对时才靠谱
- 提供专家参考答案可显著提升评判准确性
大语言模型(LLMs)的可靠评估至关重要,尤其在金融等高风险领域。现有基于LLM作为评判者的框架虽具成本低、可扩展性强的优点,但其在强调准确性的领域表现尚不明确。为此,我们构建了金融专业人员撰写的BFF-Bench数据集,包含160个难题及长文本回答,并由专家对1200条由多种LLMs生成的回答进行正确性标注(VERDICTS)。分析发现,尽管LLM评判者优于其他自动评分方法,但其与人类专家的一致性仅在模型自身能正确回答问题时成立。当未提供正确参考答案时,该一致性显著下降;而引入专家编写的标准答案后,性能大幅改善,揭示了无真人验证的LLM评判存在根本局限。
原文摘要 · Abstract (English)
Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as business and finance. The LLM-as-a-Judge framework, which uses prompted LLMs to evaluate response quality, is appealing due to its scalability, low cost, and strong correlations with human stylistic preferences. However, it remains unclear how accurately these methods can assess response quality in domains where correctness matters more than style. To address this gap, we introduce the Business and Finance Fundamentals Benchmark (BFF-Bench), a dataset of 160 challenging questions and long-form responses authored by financial professionals. These experts subsequently evaluated the correctness of 1,200 responses generated by a diverse set of LLMs on both BFF-Bench and a challenging subset of MT-Bench. With this expert-annotated dataset of judgments (VERDICTS), we analyze the agreement between a suite of automated grading methods and human experts. While we observe that LLM Judges are more reliable than other grading methods, our findings reveal a clear pattern in LLM Judge performance: when not provided with a correct reference, judges show high agreement with human experts only on questions the judges were able to correctly answer themselves. We demonstrate that providing the judges with expert-written references largely mitigates this issue, highlighting the limits of using LLM-as-a-Judge without any form of human verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。