arXiv:2502.06065cs.CLcs.AI2025-02被引 71

测试大模型对提示词微小变化的敏感度,发现现有方法效果不佳。

Benchmarking Prompt Sensitivity in Large Language Models

  • 构建新任务与数据集PromptSET,研究提示词微调对模型影响。
  • 在TriviaQA和HotpotQA上验证,多种方法预测敏感度表现有限。
  • 提醒用户:提问方式直接影响大模型回答准确性,需谨慎设计提示。

大型语言模型(LLMs)对提示词的表述变化极为敏感,可能显著影响其生成准确回复的能力。本文提出一项新任务——提示敏感度预测,并构建了名为PromptSET的数据集,用于研究细微提示变化对LLM性能的影响。基于TriviaQA和HotpotQA数据集,我们生成多种提示变体,并在多个LLM上评估其有效性。采用来自相关任务的先进方法进行基准测试,包括基于LLM的自评估、文本分类及查询性能预测技术。结果表明,现有方法难以有效解决提示敏感度预测问题,凸显出理解信息需求如何表达以获得准确响应的重要性。

原文摘要 · Abstract (English)

Large language Models (LLMs) are highly sensitive to variations in prompt formulation, which can significantly impact their ability to generate accurate responses. In this paper, we introduce a new task, Prompt Sensitivity Prediction, and a dataset PromptSET designed to investigate the effects of slight prompt variations on LLM performance. Using TriviaQA and HotpotQA datasets as the foundation of our work, we generate prompt variations and evaluate their effectiveness across multiple LLMs. We benchmark the prompt sensitivity prediction task employing state-of-the-art methods from related tasks, including LLM-based self-evaluation, text classification, and query performance prediction techniques. Our findings reveal that existing methods struggle to effectively address prompt sensitivity prediction, underscoring the need to understand how information needs should be phrased for accurate LLM responses.

大模型提示工程评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。