arXiv:2506.06950cs.CL2025-06ACL被引 19

提出评估自然语言提示质量的六维框架,揭示高效提示的关键特性。

What Makes a Good Natural Language Prompt?

  • 构建包含21项属性的六维评估框架,聚焦人类中心设计。
  • 发现单属性优化在推理任务中效果最显著,提升模型表现。
  • 指令微调增强提示可训练出更优推理模型,适合提示工程研究者。

随着大语言模型向更类人化方向发展,提示工程已成为人机交互的核心环节。然而,关于自然语言提示的质量标准尚无统一认知。本文通过元分析整合2022至2025年顶级NLP与AI会议中150余篇相关论文及博客,提出一个以属性和人为中心的提示质量评估框架,涵盖6个维度下的21项属性。研究发现现有方法对模型和任务的支持不均衡,存在显著研究空白。进一步分析高质量提示中属性间的相关性,提炼出实用提示建议。实证研究表明,在推理任务中,单一属性增强往往带来最大收益。最后,基于增强提示进行指令微调,可训练出性能更优的推理模型。该工作为属性导向的提示评估与优化奠定基础,推动人机沟通与提示研究的发展。

原文摘要 · Abstract (English)

As large language models (LLMs) have progressed towards more human-like and human--AI communications have become prevalent, prompting has emerged as a decisive component. However, there is limited conceptual consensus on what exactly quantifies natural language prompts. We attempt to address this question by conducting a meta-analysis surveying more than 150 prompting-related papers from leading NLP and AI conferences from 2022 to 2025 and blogs. We propose a property- and human-centric framework for evaluating prompt quality, encompassing 21 properties categorized into six dimensions. We then examine how existing studies assess their impact on LLMs, revealing their imbalanced support across models and tasks, and substantial research gaps. Further, we analyze correlations among properties in high-quality natural language prompts, deriving prompting recommendations. We then empirically explore multi-property prompt enhancements in reasoning tasks, observing that single-property enhancements often have the greatest impact. Finally, we discover that instruction-tuning on property-enhanced prompts can result in better reasoning models. Our findings establish a foundation for property-centric prompt evaluation and optimization, bridging the gaps between human--AI communication and opening new prompting research directions.

提示工程大模型评估框架人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。