arXiv:2511.10661cs.CLcs.LG2025-11中稿 · NeurIPS被引 2

用贝叶斯方法量化大模型生成行为的不确定性,更可靠评估其安全表现。

Bayesian Evaluation of Large Language Model Behavior

  • 采用贝叶斯框架分析大模型输出的二元评价结果,考虑生成过程的随机性。
  • 在对抗输入和对话偏好测试中,准确给出拒绝率与模型优劣的置信区间。
  • 适合关注模型评估可靠性、需量化不确定性的研究人员和工程团队。

随着基于大语言模型(LLM)的文本生成系统广泛应用,评估其行为(如产生有害内容或对对抗输入敏感)变得愈发重要。现有评估多依赖人工标注的提示集,通过二元判断(如有害/非有害)聚合得分。然而,这些方法常忽略统计不确定性。本文面向应用统计学读者,介绍大模型文本生成与评估背景,并提出一种贝叶斯方法来量化二元评价指标的不确定性,尤其关注由模型概率生成策略带来的随机性影响。通过两个案例验证:1)在用于诱使有害响应的对抗输入基准上评估拒绝率;2)在开放式对话样本上评估一个模型对另一个模型的成对偏好。结果表明,该方法能有效提供对大模型行为的不确定性量化,增强评估可信度。

原文摘要 · Abstract (English)

It is increasingly important to evaluate how text generation systems based on large language models (LLMs) behave, such as their tendency to produce harmful output or their sensitivity to adversarial inputs. Such evaluations often rely on a curated benchmark set of input prompts provided to the LLM, where the output for each prompt may be assessed in a binary fashion (e.g., harmful/non-harmful or does not leak/leaks sensitive information), and the aggregation of binary scores is used to evaluate the LLM. However, existing approaches to evaluation often neglect statistical uncertainty quantification. With an applied statistics audience in mind, we provide background on LLM text generation and evaluation, and then describe a Bayesian approach for quantifying uncertainty in binary evaluation metrics. We focus in particular on uncertainty that is induced by the probabilistic text generation strategies typically deployed in LLM-based systems. We present two case studies applying this approach: 1) evaluating refusal rates on a benchmark of adversarial inputs designed to elicit harmful responses, and 2) evaluating pairwise preferences of one LLM over another on a benchmark of open-ended interactive dialogue examples. We demonstrate how the Bayesian approach can provide useful uncertainty quantification about the behavior of LLM-based systems.

大模型评估贝叶斯方法不确定性量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。