用分层模糊概率模型评估大模型可靠性,更真实反映实际使用中的表现。
A Hierarchical Imprecise Probability Approach to Reliability Assessment of Large Language Models
- 构建分层模糊概率框架,从子领域到系统级推断可靠性
- 通过先验不确定性和使用场景建模,给出可靠性置信区间
- 适合关注大模型部署安全性的研究者与工程团队
大语言模型(LLMs)在多个领域广泛应用,亟需严谨的可靠性评估方法。现有基于基准的评估主要提供模型在数据集上的准确率统计,难以揭示模型在真实运行环境下的概率行为。本文提出HIP-LLM,一种分层模糊概率框架,用于建模和推断LLM可靠性。基于软件可靠性工程,将可靠性定义为在给定运行剖面下完成指定任务数时无故障运行的概率。该框架分层表示跨(子)领域的依赖关系,实现从子域到系统级的多层级推理;引入模糊先验以捕捉认知不确定性,并结合运行剖面反映实际使用情境。通过后验推断获得可靠性置信区间,量化先验与数据共同带来的不确定性。在多个基准数据集上的实验表明,相比现有基准方法与先进方法,HIP-LLM能提供更准确、标准化的可靠性表征。相关代码已公开。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly deployed across diverse domains, raising the need for rigorous reliability assessment methods. Existing benchmark-based evaluations primarily offer descriptive statistics of model accuracy over datasets, providing limited insight into the probabilistic behavior of LLMs under real operational conditions. This paper introduces HIP-LLM, a Hierarchical Imprecise Probability framework for modeling and inferring LLM reliability. Building upon the foundations of software reliability engineering, HIP-LLM defines LLM reliability as the probability of failure-free operation over a specified number of future tasks under a given Operational Profile (OP). HIP-LLM represents dependencies across (sub-)domains hierarchically, enabling multi-level inference from subdomain to system-level reliability. HIP-LLM embeds imprecise priors to capture epistemic uncertainty and incorporates OPs to reflect usage contexts. It derives posterior reliability envelopes that quantify uncertainty across priors and data. Experiments on multiple benchmark datasets demonstrate that HIP-LLM offers a more accurate and standardized reliability characterization than existing benchmark and state-of-the-art approaches. A publicly accessible repository of HIP-LLM is provided.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。