arXiv:2608.18539cs.LGcs.AI2026-08中稿 · ICML被引 2

用交互分析法揭示大模型提示敏感性的内在机制

Evaluating and Explaining Prompt Sensitivity of LLMs Using Interactions

论文配图:Evaluating and Explaining Prompt Sensitivity of LLMs Using Interactions
图 1 · 摘自论文原文
  • 通过分解输出得分中的交互项,捕捉提示细微变化的影响
  • 发现即使输出不变,交互项仍可能剧烈波动,揭示敏感性根源
  • 适合关注模型鲁棒性与可解释性的研究者参考

大型语言模型(LLMs)虽具备强大能力,但其性能易受提示扰动影响,即提示敏感性。现有评估方法仅比较输出差异,无法解释内部原因。本文引入交互分析作为细粒度工具,将模型输出得分分解为多组交互项,每项反映输入变量间的非线性关系。实验发现,微小提示变化可导致交互项严重不稳定,即便最终输出保持一致。为此,提出基于交互的提示敏感性(IPS)度量方法,量化提示微调时交互项的变化。在50个开源模型上应用该度量,识别出四种降低敏感性的因素:监督微调、模型规模增大、密集架构、少样本学习。更关键的是,这四类因素均通过抑制低阶交互(涉及较少输入变量)的敏感性来提升稳定性。

原文摘要 · Abstract (English)

The remarkable capabilities of large language models (LLMs) are often undermined by their instability. Even subtle and semantically irrelevant changes in prompts can cause dramatic fluctuations in performance, a phenomenon known as prompt sensitivity. Previous studies typically evaluate prompt sensitivity by comparing the LLM's final outputs when prompts change. However, such coarse-grained metrics fail to explain the internal reasons for prompt sensitivity. In this paper, we introduce interactions as a fine-grained tool to analyze prompt sensitivity of LLMs. Specifically, we decompose the output score of the LLM into a set of interactions. Each interaction represents a nonlinear relationship involving a set of input variables. We discover that subtle changes to prompts can trigger severe instability in interactions, even when the outputs of the LLM remain the same. To this end, we propose an Interaction-based Prompt Sensitivity (IPS) metric by quantifying changes in interactions when we introduce subtle changes to prompts. We apply the IPS metric to 50 open-source LLMs and uncover four factors that reduce the prompt sensitivity of LLMs, including supervised fine-tuning, increased model scales, dense architectures, and few-shot learning. More crucially, we discover a common mechanism by which these four factors reduce prompt sensitivity: all four factors tend to reduce the prompt sensitivity of low-order interactions (i.e., interactions involving few input variables).

提示敏感性交互分析大模型鲁棒性可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。