用检查清单式提示词,能显著提升大模型回答质量并减少用户反复沟通。
Less Back-and-Forth: A Comparative Study of Structured Prompting

- 用检查清单优化提示词,引导模型更完整地响应任务。
- 检查清单提示使评分平均达7.50分(满分8分),优于原始提示的5.67分。
- 适合希望高效获得高质量回复的用户,尤其在写作、编程等场景中。
大语言模型广泛应用于开放性任务,但模糊的提示常导致回答质量低且需多次交互。本文比较了三种提示设计:原始提示、检查清单优化提示和澄清问题提示。在摘要生成、规划、解释和编程四类任务上,使用ChatGPT、Claude和Grok三类模型进行评估。所有输出均采用统一评分标准,涵盖任务完成度、正确性、合规性和清晰度。检查清单提示得分最高,平均7.50分(满分8分),显著优于原始提示的5.67分和澄清问题提示的6.67分。同时,检查清单提示使用的平均token数更少,表现出更优的质量-效率平衡。结果表明,简单的提示清单可有效提升模型输出质量并减少用户额外交互。
原文摘要 · Abstract (English)
Large language models (LLMs) are widely used for open-ended tasks, but underspecified prompts can lead to low-quality answers and additional interaction. This paper studies whether structured prompt design improves response quality while reducing user effort. We compare three prompt conditions: a raw prompt, a checklist-improved prompt, and a clarifying-question prompt. We evaluate these conditions across four task types--summarization, planning, explanation, and coding--using three LLM systems: ChatGPT, Claude, and Grok. Each output is scored with a unified rubric covering task completion, correctness, compliance, and clarity. Checklist-improved prompts achieved the highest mean rubric score, 7.50 out of 8, compared with 5.67 for raw prompts and 6.67 for clarifying-question prompts. Checklist prompts also produced the best quality-effort tradeoff, using fewer average tokens than both raw and clarifying prompts. These results suggest that a simple prompt checklist can improve LLM responses while reducing unnecessary interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。