为大模型评估指标设定可靠阈值,避免部署风险。
How to Choose a Threshold for an Evaluation Metric for Large Language Models
- 基于应用风险和利益相关者容忍度,分步确定阈值。
- 使用真实数据和统计方法,确保阈值科学合理。
- 适用于大模型及通用生成式AI的阈值选择框架。
为确保大语言模型(LLMs)的可靠性和可监控性,学术界提出了多种评估指标。然而,关于如何为这些指标设定稳健阈值的研究却很少,而阈值选择不当在模型部署中可能带来严重后果。借鉴金融等受监管行业的传统模型风险管理(MRM)指南,本文提出一个逐步实施的阈值选择方法。该方法强调首先识别特定应用的风险及利益相关者的风险容忍度,随后利用可用的真实数据,采用严谨的统计程序确定给定评估指标的阈值。以公开的HaluBench数据集和多个开源库中的忠实性(Faithfulness)指标为例,展示了该方法的实际应用。研究还为构建系统性阈值选择方法奠定了基础,不仅适用于大模型,也适用于其他生成式AI应用。
原文摘要 · Abstract (English)
To ensure and monitor large language models (LLMs) reliably, various evaluation metrics have been proposed in the literature. However, there is little research on prescribing a methodology to identify a robust threshold on these metrics even though there are many serious implications of an incorrect choice of the thresholds during deployment of the LLMs. Translating the traditional model risk management (MRM) guidelines within regulated industries such as the financial industry, we propose a step-by-step recipe for picking a threshold for a given LLM evaluation metric. We emphasize that such a methodology should start with identifying the risks of the LLM application under consideration and risk tolerance of the stakeholders. We then propose concrete and statistically rigorous procedures to determine a threshold for the given LLM evaluation metric using available ground-truth data. As a concrete example to demonstrate the proposed methodology at work, we employ it on the Faithfulness metric, as implemented in various publicly available libraries, using the publicly available HaluBench dataset. We also lay a foundation for creating systematic approaches to select thresholds, not only for LLMs but for any GenAI applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。