arXiv:2412.00868cs.LGcs.CL2024-12被引 9

提出新方法量化大模型输入扰动的影响,解决输出随机性干扰问题。

Quantifying perturbation impacts for large language models

  • 将扰动分析转为频率学假设检验,用蒙特卡洛采样构建语义空间分布。
  • 在低维语义空间中比较输出分布,可得可解释的p值和效应量。
  • 适用于任意黑盒大模型,支持多种扰动测试且控制错误率。

本文研究如何量化输入扰动对大语言模型(LLM)输出的影响,这是评估模型可靠性与事后可解释性的基础任务。该领域的主要挑战在于区分模型响应中的有意义变化与固有的输出随机性。为此,我们提出分布基扰动分析(DBPA),将LLM扰动分析重新建模为频率学假设检验问题。DBPA通过蒙特卡洛采样,在低维语义相似性空间中构建经验零假设和备择假设输出分布。在降维空间中比较蒙特卡洛估计,实现无需强分布假设的可计算频率推断。该框架具有模型无关性,支持对任意输入扰动在任何黑盒LLM上的评估,输出可解释的p值,支持多扰动测试并控制误差率,且为任意相似性或距离度量提供标量效应量。我们验证了DBPA在评估扰动影响方面的有效性,展示了其在扰动分析中的通用性。

原文摘要 · Abstract (English)

We consider the problem of quantifying how an input perturbation impacts the outputs of large language models (LLMs), a fundamental task for model reliability and post-hoc interpretability. A key obstacle in this domain is disentangling the meaningful changes in model responses from the intrinsic stochasticity of LLM outputs. To overcome this, we introduce Distribution-Based Perturbation Analysis (DBPA), a framework that reformulates LLM perturbation analysis as a frequentist hypothesis testing problem. DBPA constructs empirical null and alternative output distributions within a low-dimensional semantic similarity space via Monte Carlo sampling. Comparisons of Monte Carlo estimates in the reduced dimensionality space enables tractable frequentist inference without relying on restrictive distributional assumptions. The framework is model-agnostic, supports the evaluation of arbitrary input perturbations on any black-box LLM, yields interpretable p-values, supports multiple perturbation testing via controlled error rates, and provides scalar effect sizes for any chosen similarity or distance metric. We demonstrate the effectiveness of DBPA in evaluating perturbation impacts, showing its versatility for perturbation analysis.

大模型扰动分析可解释性统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。