arXiv:2410.14361cs.CL2024-10EMNLP

提出快速计算语言模型对上下文敏感度的新方法。

Efficiently Computing Susceptibility to Context in Language Models

  • 基于费舍尔信息设计高效估算方法
  • 速度比传统方法快70倍且结果相近
  • 适用于分析模型敏感度影响因素

现代语言模型在回答问题时能融入用户输入的上下文信息,但对上下文细微变化的敏感度不一。为量化这种敏感性,Du等(2024)提出了信息论指标‘易感性’(susceptibility),衡量上下文对模型响应分布的影响程度。然而精确计算该指标困难,现有方法依赖蒙特卡洛近似,需大量采样,效率低下。本文提出‘费舍尔易感性’(Fisher susceptibility),基于费舍尔信息实现高效估计。实验表明,在多种查询领域中,该方法与蒙特卡洛估计结果相当,同时提速70倍。利用其高效率,我们进一步分析了影响语言模型易感性的因素,发现大模型与小模型具有相似的易感性水平。

原文摘要 · Abstract (English)

One strength of modern language models is their ability to incorporate information from a user-input context when answering queries. However, they are not equally sensitive to the subtle changes to that context. To quantify this, Du et al. (2024) gives an information-theoretic metric to measure such sensitivity. Their metric, susceptibility, is defined as the degree to which contexts can influence a model's response to a query at a distributional level. However, exactly computing susceptibility is difficult and, thus, Du et al. (2024) falls back on a Monte Carlo approximation. Due to the large number of samples required, the Monte Carlo approximation is inefficient in practice. As a faster alternative, we propose Fisher susceptibility, an efficient method to estimate the susceptibility based on Fisher information. Empirically, we validate that Fisher susceptibility is comparable to Monte Carlo estimated susceptibility across a diverse set of query domains despite its being $70\times$ faster. Exploiting the improved efficiency, we apply Fisher susceptibility to analyze factors affecting the susceptibility of language models. We observe that larger models are as susceptible as smaller ones.

语言模型易感性效率费舍尔信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。