不同语气提示会显著影响大模型回答质量和推理成本。
Understanding Tone-Dependent Inference Cost in Large Language Models
- 通过七种语气测试模型输出长度与准确率的权衡
- 语气变化导致输出词元量最高相差44.3%
- 粗鲁语气在部分模型中更高效,适合资源敏感场景
我们研究了提示语气如何影响大语言模型的回答准确率和推理成本(以输出词元消耗衡量)。在包含570个问题的MMLU数据集上,对七种从奉承到威胁的不同语气进行了实验。结果表明,所有模型的输出词元长度变化远超准确率变化,最大差异达44.3%。我们还分析了答案准确率与推理过程平均输出词元长度之间的权衡关系:ChatGPT 4o和5-nano模型中,粗鲁语气表现最优;Gemini 2.5 Flash和2.5 Flash Lite模型中,粗鲁与中性语气位于帕累托最优前沿。研究发现,提示语气不仅影响回答质量,还显著影响现代大模型的计费推理资源消耗。
原文摘要 · Abstract (English)
We examine how prompt tone affects both accuracy of the LLM answers and inference cost as reflected in output-token consumption. Experiments were performed to understand the trade-offs between accuracy and inference cost on a 570 Question MMLU dataset for LLM models prompted in seven different tones from sycophantic to threatening. Our results show that the output-token-length variation substantially exceeded accuracy variation across all models. Output-token consumption varied by up to 44.3% across tone conditions. We also analyzed the tradeoff between the accuracy of the answers and the average output token length in the reasoning process. For the ChatGPT models 4o and 5-nano, the rude tone is quite dominant. For the Gemini models 2.5 Flash and 2.5 Flash Lite, the rude and neutral tones are dominant on the Pareto-optimal frontier. We find that prompt tone influences not only answer quality but also the amount of billable inference resources consumed by modern LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。