研究提示词详细程度如何影响大模型推理,发现越具体越准确。
DETAIL Matters: Measuring the Impact of Prompt Specificity on Reasoning in Large Language Models
- 用困惑度量化提示词详细程度,生成多级提示
- 小模型和流程类任务中,详细提示可提升准确率
- 为模型适配提示提供工具与数据,适合提示工程研究者
提示词设计对大语言模型的推理表现至关重要,但提示词的具体程度(即详细或模糊)的影响仍缺乏研究。本文提出DETAIL框架,通过GPT-4生成多层级提示,利用困惑度量化提示具体程度,并采用基于GPT的语义等价性评估正确性。在30个新推理任务上对GPT-4和O3-mini进行实验,结果表明,提示词具体性可提升准确性,尤其对较小模型和程序类任务效果更显著。研究强调了自适应提示策略的重要性,并提供了工具与数据集以支持后续研究。
原文摘要 · Abstract (English)
Prompt design plays a critical role in the reasoning performance of large language models (LLMs), yet the impact of prompt specificity - how detailed or vague a prompt is - remains understudied. This paper introduces DETAIL, a framework for evaluating LLM performance across varying levels of prompt specificity. We generate multi-level prompts using GPT-4, quantify specificity via perplexity, and assess correctness using GPT-based semantic equivalence. Experiments on 30 novel reasoning tasks across GPT-4 and O3-mini reveal that specificity improves accuracy, especially for smaller models and procedural tasks. Our results highlight the need for adaptive prompting strategies and provide tools and data to support further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。