arXiv:2410.12405cs.CL2024-10EMNLP被引 226

提出新方法评估大模型对提示的敏感性,发现大模型更抗干扰。

ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs

  • 设计新指标PromptSensiScore和利用解码置信度分析敏感机制。
  • 大模型比小模型更抗提示变化,少量示例可降低敏感性。
  • 适合研究提示工程、模型评估与主观评测的科研人员使用。

大型语言模型(LLMs)在多种任务中表现优异,但其性能高度依赖提示词。这种波动给准确评估和用户满意度带来挑战。现有研究常忽视实例级提示变化及其对主观评价的影响。为此,我们提出ProSA框架,用于评估和理解大模型的提示敏感性。ProSA引入新颖的敏感性指标PromptSensiScore,并利用解码置信度揭示内在机制。我们在多个任务上的广泛研究发现,提示敏感性在不同数据集和模型间存在差异,大模型表现出更强的鲁棒性。少量示例能缓解敏感问题,且主观评价在复杂推理任务中尤其受提示影响。此外,模型置信度越高,提示鲁棒性越强。本工作为研究大模型提示敏感性提供了有力工具。项目已开源:https://github.com/open-compass/ProSA。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated impressive capabilities across various tasks, but their performance is highly sensitive to the prompts utilized. This variability poses challenges for accurate assessment and user satisfaction. Current research frequently overlooks instance-level prompt variations and their implications on subjective evaluations. To address these shortcomings, we introduce ProSA, a framework designed to evaluate and comprehend prompt sensitivity in LLMs. ProSA incorporates a novel sensitivity metric, PromptSensiScore, and leverages decoding confidence to elucidate underlying mechanisms. Our extensive study, spanning multiple tasks, uncovers that prompt sensitivity fluctuates across datasets and models, with larger models exhibiting enhanced robustness. We observe that few-shot examples can alleviate this sensitivity issue, and subjective evaluations are also susceptible to prompt sensitivities, particularly in complex, reasoning-oriented tasks. Furthermore, our findings indicate that higher model confidence correlates with increased prompt robustness. We believe this work will serve as a helpful tool in studying prompt sensitivity of LLMs. The project is released at: https://github.com/open-compass/ProSA .

大模型提示敏感性评估方法鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。