研究提示语礼貌程度对大模型准确率的影响,发现越不客气反而表现更好。
Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy (short paper)
- 设计五种语气的提示语,对比模型在250个问题上的表现
- 不客气提示准确率达84.8%,比最客气的80.8%更高
- 揭示新模型对语气敏感性与以往认知不同,适合提示工程研究者
自然语言提示的表述方式已被证明会影响大语言模型(LLMs)的表现,但礼貌和语气的作用仍缺乏研究。本研究探讨不同礼貌程度的提示语如何影响模型在多项选择题上的准确率。我们构建了一个包含50个基础问题的数据集,覆盖数学、科学和历史,每个问题改写为五种语气变体:非常礼貌、礼貌、中性、粗鲁、非常粗鲁,共生成250个独特提示。使用ChatGPT 4o评估各条件下的回答,并采用配对样本t检验分析统计显著性。结果出乎意料:不礼貌提示始终优于礼貌提示,准确率从非常礼貌的80.8%提升至非常粗鲁的84.8%。该发现与早期认为粗鲁导致表现下降的研究相悖,表明新一代大模型可能对语气变化有不同反应。研究强调了提示语实用层面的重要性,并引发关于人机交互社会维度的更广泛思考。
原文摘要 · Abstract (English)
The wording of natural language prompts has been shown to influence the performance of large language models (LLMs), yet the role of politeness and tone remains underexplored. In this study, we investigate how varying levels of prompt politeness affect model accuracy on multiple-choice questions. We created a dataset of 50 base questions spanning mathematics, science, and history, each rewritten into five tone variants: Very Polite, Polite, Neutral, Rude, and Very Rude, yielding 250 unique prompts. Using ChatGPT 4o, we evaluated responses across these conditions and applied paired sample t-tests to assess statistical significance. Contrary to expectations, impolite prompts consistently outperformed polite ones, with accuracy ranging from 80.8% for Very Polite prompts to 84.8% for Very Rude prompts. These findings differ from earlier studies that associated rudeness with poorer outcomes, suggesting that newer LLMs may respond differently to tonal variation. Our results highlight the importance of studying pragmatic aspects of prompting and raise broader questions about the social dimensions of human-AI interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。