重复提问能提升大模型准确率,但效果不显著。
Asking Again and Again: Exploring LLM Robustness to Repeated Questions
- 在提示中重复提问,观察模型注意力聚焦变化。
- 最高可提升6%准确率,但未达统计显著水平。
- 适合研究提示工程与模型鲁棒性的读者。
本研究探究在提示中重复提问是否影响大语言模型(LLMs)的表现。我们假设在同一提示中重述问题可能增强模型对查询关键要素的关注。我们在三个阅读理解数据集上,评估了五种近期LLM(包括GPT-4o-mini、DeepSeek-V3及小型开源模型)在不同提示设置下的表现,其中问题重复次数为1、3或5次。结果表明,问题重复可使模型准确率最高提升6%。然而,在所有模型、设置和数据集上,该提升均未达到统计显著性。研究为提示设计与大模型行为提供了新见解,表明仅靠重复提问无法显著改善输出质量。
原文摘要 · Abstract (English)
This study investigates whether repeating questions within prompts influences the performance of large language models (LLMs). We hypothesize that reiterating a question within a single prompt might enhance the model's focus on key elements of the query. We evaluate five recent LLMs -- including GPT-4o-mini, DeepSeek-V3, and smaller open-source models -- on three reading comprehension datasets under different prompt settings, varying question repetition levels (1, 3, or 5 times per prompt). Our results demonstrate that question repetition can increase models' accuracy by up to $6\%$. However, across all models, settings, and datasets, we do not find the result statistically significant. These findings provide insights into prompt design and LLM behavior, suggesting that repetition alone does not significantly impact output quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。