不同语气提示会影响大模型答题准确率,且影响因模型而异。
Mind Your Tone: Does Tone Alter LLM Performance?
- 测试多种语气提示对模型表现的影响
- 部分模型在不同语气下准确率差异超10%
- 适合关注提示工程与模型稳定性的研究者
大型语言模型(LLM)应用日益广泛,但其性能受提示风格和语气影响。本研究探究了提示语气变化是否以及如何影响模型在客观多选题上的准确性。使用两个数据集:包含50道题的5种语气变体数据集,以及涵盖57个学科的570道题的MMLU子集,共7种语气变体。评估了四种低成本、流行的模型:ChatGPT-4o、ChatGPT-5-nano、Gemini 2.5 Flash和Gemini 2.5 Flash Lite。结果显示,语气影响具有系统性,但高度依赖具体模型:某些模型仅出现微小但统计显著的误差变化,另一些则在不同语气下准确率波动超过10%。此外,不同学科对语气敏感度存在差异,并提出一种路由框架以解释语气如何调节内部推理模式。研究警示用户,不应假设模型在语气变化下仍具鲁棒性。
原文摘要 · Abstract (English)
The use of Large Language Models (LLMs) is proliferating, yet their performance is observed to vary based on prompting styles and tones. In this study, we investigate both whether and how tonal variations in prompts lead to disparate LLM accuracy for objective multiple-choice questions. We use two datasets: a 50-base question dataset with five tone variants and a 570-base question MMLU subset spanning 57 subjects with seven tone variants. Experiments were conducted to evaluate the performance of four cost-efficient, popular LLMs: ChatGPT-4o, ChatGPT-5-nano, Gemini 2.5 Flash, and Gemini 2.5 Flash Lite. Across models, tonal effects are systematic but highly model-dependent. Some models show small, yet statistically significant, shifts, while others exhibit large accuracy swings across tones. Further, we identify subject-level differences in tone sensitivity and present a routing framework to explain how tones may attune internal reasoning modes. Our findings caution users against assuming tone-robust reliability in LLM deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。