arXiv:2506.07947cs.CL2025-06ICML被引 6

用统计检验方法量化大模型响应变化,判断干预是否有效。

Statistical Hypothesis Testing for Auditing Robustness in Language Models

  • 将模型输出变化转为频率学假设检验问题,基于语义空间采样构建分布。
  • 可检测任意输入扰动下的响应变化,给出可解释的p值与效应量。
  • 适合做模型鲁棒性审计,尤其适用于黑盒模型的可靠性评估。

当大语言模型(LLM)系统在输入扰动或模型变体切换下输出发生变化时,如何判断这种变化是否显著?由于模型输出具有随机性,直接比较输出或整体分布均不可行。现有方法多聚焦于偏见或公平性分析,不适用于此类问题。为此,本文提出基于分布的扰动分析框架,将扰动分析重构为频率学假设检验问题。通过蒙特卡洛采样,在低维语义相似性空间中构建经验零假设与备择假设分布,实现无需强分布假设的可计算推断。该框架具备五项优势:(i) 模型无关;(ii) 支持任意输入扰动对任意黑盒模型的评估;(iii) 输出可解释的p值;(iv) 可控误差率支持多重扰动测试;(v) 提供标量效应大小。在多个案例研究中验证了其有效性,可量化响应变化、测量真/假阳性率,并评估与参考模型的一致性。整体而言,该框架为大模型审计提供了可靠的频率学检验基础。

原文摘要 · Abstract (English)

Consider the problem of testing whether the outputs of a large language model (LLM) system change under an arbitrary intervention, such as an input perturbation or changing the model variant. We cannot simply compare two LLM outputs since they might differ due to the stochastic nature of the system, nor can we compare the entire output distribution due to computational intractability. While existing methods for analyzing text-based outputs exist, they focus on fundamentally different problems, such as measuring bias or fairness. To this end, we introduce distribution-based perturbation analysis, a framework that reformulates LLM perturbation analysis as a frequentist hypothesis testing problem. We construct empirical null and alternative output distributions within a low-dimensional semantic similarity space via Monte Carlo sampling, enabling tractable inference without restrictive distributional assumptions. The framework is (i) model-agnostic, (ii) supports the evaluation of arbitrary input perturbations on any black-box LLM, (iii) yields interpretable p-values; (iv) supports multiple perturbations via controlled error rates; and (v) provides scalar effect sizes. We demonstrate the usefulness of the framework across multiple case studies, showing how we can quantify response changes, measure true/false positive rates, and evaluate alignment with reference models. Above all, we see this as a reliable frequentist hypothesis testing framework for LLM auditing.

模型审计统计检验鲁棒性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。