用统计方法评估大模型判断的可靠性,发现单次输出不可靠。
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
- 引入麦当劳奥米茄系数评估模型判断的一致性
- 多样本测试显示温度影响判断可靠性,单次结果易偏差
- 提醒避免依赖单一输出,适合开发可信AI系统的团队
大语言模型(LLMs)虽强大且广泛应用,但其随机性影响输出可靠性。即使在确定性设置下,单次采样仍可能误导。本文基于LLM-as-a-judge思想,提出新框架,利用麦当劳奥米茄系数严格评估LLM在标准单轮与多轮基准上对其他LLM输出的判断可靠性,并同时研究温度对可靠性的影响。分析表明,固定随机性无法保证可靠,必须考虑多样本;这对手下游应用有重要启示。研究强调需深入理解LLM可靠性,警惕过度依赖单次评估的风险。本工作为构建更可信的LLM系统迈出关键一步。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have become increasingly powerful and ubiquitous, but their stochastic nature poses challenges to the reliability of their outputs. While deterministic settings can improve consistency, they do not guarantee reliability, as a single sample from the model's probability distribution can still be misleading. Building upon the concept of LLM-as-a-judge, we introduce a novel framework for rigorously evaluating the reliability of LLM judgments, leveraging McDonald's omega. We evaluate the reliability of LLMs when judging the outputs of other LLMs on standard single-turn and multi-turn benchmarks, simultaneously investigating the impact of temperature on reliability. By analyzing these results, we demonstrate the limitations of fixed randomness and the importance of considering multiple samples, which we show has significant implications for downstream applications. Our findings highlight the need for a nuanced understanding of LLM reliability and the potential risks associated with over-reliance on single-shot evaluations. This work provides a crucial step towards building more trustworthy and reliable LLM-based systems and applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。