提出统一评估开源大模型可靠性的综合得分方法。
Beyond Hallucinations: A Composite Score for Measuring Reliability in Open-Source Large Language Models
- 设计综合可靠性评分(CRS),融合校准、鲁棒性与不确定性量化。
- 在5个数据集上测试10个模型,发现评分稳定且能暴露单一指标忽略的缺陷。
- 适合关注模型可信度的医疗、金融等领域研究者使用。
LLaMA、Mistral、Gemma等大型语言模型在医疗、法律、金融等关键领域应用日益广泛,但其可靠性仍不明确。这些模型常出现过度自信的错误,在输入变化下性能下降,且缺乏清晰的不确定性估计。现有评估方法分散,仅关注孤立方面。本文提出综合可靠性评分(CRS),将校准性、鲁棒性和不确定性量化整合为单一可解释指标。在五个问答数据集上对十款主流开源LLM进行基准、扰动和校准方法测试,结果表明:CRS保持稳定的模型排名,揭示了单一指标遗漏的隐藏失效模式,并指出最可靠的系统需在准确率、鲁棒性与校准不确定性之间取得平衡。
原文摘要 · Abstract (English)
Large Language Models (LLMs) like LLaMA, Mistral, and Gemma are increasingly used in decision-critical domains such as healthcare, law, and finance, yet their reliability remains uncertain. They often make overconfident errors, degrade under input shifts, and lack clear uncertainty estimates. Existing evaluations are fragmented, addressing only isolated aspects. We introduce the Composite Reliability Score (CRS), a unified framework that integrates calibration, robustness, and uncertainty quantification into a single interpretable metric. Through experiments on ten leading open-source LLMs across five QA datasets, we assess performance under baselines, perturbations, and calibration methods. CRS delivers stable model rankings, uncovers hidden failure modes missed by single metrics, and highlights that the most dependable systems balance accuracy, robustness, and calibrated uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。