用隐马尔可夫模型捕捉大模型响应的依赖性,更准确评估可靠性。
Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model

- 引入隐马尔可夫模型建模大模型响应间的序列依赖关系。
- 实验证明忽略依赖会高估可靠性,误差可达15%以上。
- 适合关注模型长期行为与可信度评估的研究者。
大语言模型(LLMs)的可靠性评估旨在估计模型在特定使用场景下产生正确回答的概率。传统基于基准的评估通常以整体准确率作为点估计,但未能刻画可靠性结论相关的不确定性。目前,统计推断方法正用于评估LLM可靠性,但其关键假设是测试结果可视为独立重复试验。这一假设在序列化场景中可能不成立,因为后续响应可能受先前交互中保留的上下文、错误传播或动态交互状态的影响。本文通过放宽任务结果独立性的假设,将隐马尔可夫模型(HMM)引入分层贝叶斯框架,以捕捉由基准构建的交互会话中的序列依赖性。在此设定中,输出由一个一阶马尔可夫过程驱动的潜在交互状态生成,从而反映上下文的变化。在Anthropic Claude和OpenAI模型上,针对四个数据集的实验表明,序列依赖对可靠性评估有显著影响。结果表明,忽略序列依赖可能导致可靠性估计过于自信,误差范围达15%以上。
原文摘要 · Abstract (English)
Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a specified operational profile. Conventional benchmark-based evaluation, often summarized by aggregate accuracy, provides a point estimate of performance but does not characterize the uncertainty associated with reliability claims. Currently, statistical inference methods for LLM reliability assessment are emerging. However, a key assumption underlying these models is that test outcomes can be treated as independent repeated trials. This assumption may be inappropriate in sequential settings, where later responses depend on earlier interactions through retained context, error propagation, or an evolving interaction state. We extend a hierarchical Bayesian framework for LLM reliability assessment by relaxing the assumption of independent task outcomes and introducing a Hidden Markov Model to capture sequential dependence in benchmark-constructed interaction sessions. In this formulation, outcomes are generated from a latent interaction state evolving according to a first-order Markov process, capturing changes in interaction context. Through experiments using Anthropic Claude and OpenAI on four datasets, we demonstrate the potential impact of sequential dependence on reliability assessment. The results suggest that ignoring sequential dependence may lead to overconfident reliability estimates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。