测试语气对大模型回答的影响,发现特定领域下粗鲁提问会降低准确率。
Does Tone Change the Answer? Evaluating Prompt Politeness Effects on Modern LLMs: GPT, Gemini, and LLaMA
- 设计三类语气提示(礼貌/中性/粗鲁)对比模型表现。
- 人文类任务中粗鲁语气使GPT和Llama准确率下降,Gemini不受影响。
- 跨任务汇总后效果不显著,说明模型整体对语气较鲁棒。
提示工程已成为影响大语言模型性能的关键因素,但语言语调与礼貌程度等实用要素的影响仍研究不足,尤其在不同模型家族间的差异。本文提出系统评估框架,考察交互语气对模型准确率的影响,针对GPT-4o mini(OpenAI)、Gemini 2.0 Flash(Google DeepMind)和Llama 4 Scout(Meta)三款近期广泛可用的LLM进行测试。基于MMMLU基准,我们在六项涵盖理工与人文学科的任务上,评估非常礼貌、中性和非常粗鲁提示下的模型表现,并通过统计显著性检验分析成对准确率差异。结果显示,语气敏感性具有模型依赖性和领域特异性:中性或非常礼貌提示通常优于非常粗鲁提示,但显著效应仅出现在部分人文学科任务中——粗鲁语气使GPT和Llama准确率下降,而Gemini则相对不敏感。当按领域聚合任务表现时,语气效应减弱且大多失去统计显著性。相比早期研究,这些发现表明数据集规模与覆盖范围显著影响语气效应的检测。总体而言,虽然在特定解释性场景下语气可能影响结果,但现代大模型在典型多领域使用中对语气变化普遍具有鲁棒性,为实际部署中的提示设计与模型选择提供实践指导。
原文摘要 · Abstract (English)
Prompt engineering has emerged as a critical factor influencing large language model (LLM) performance, yet the impact of pragmatic elements such as linguistic tone and politeness remains underexplored, particularly across different model families. In this work, we propose a systematic evaluation framework to examine how interaction tone affects model accuracy and apply it to three recently released and widely available LLMs: GPT-4o mini (OpenAI), Gemini 2.0 Flash (Google DeepMind), and Llama 4 Scout (Meta). Using the MMMLU benchmark, we evaluate model performance under Very Polite, Neutral, and Very Rude prompt variants across six tasks spanning STEM and Humanities domains, and analyze pairwise accuracy differences with statistical significance testing. Our results show that tone sensitivity is both model-dependent and domain-specific. Neutral or Very Polite prompts generally yield higher accuracy than Very Rude prompts, but statistically significant effects appear only in a subset of Humanities tasks, where rude tone reduces accuracy for GPT and Llama, while Gemini remains comparatively tone-insensitive. When performance is aggregated across tasks within each domain, tone effects diminish and largely lose statistical significance. Compared with earlier research, these findings suggest that dataset scale and coverage materially influence the detection of tone effects. Overall, our study indicates that while interaction tone can matter in specific interpretive settings, modern LLMs are broadly robust to tonal variation in typical mixed-domain use, providing practical guidance for prompt design and model selection in real-world deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。