提示工程效果因任务而异,无普适标准。
Prompting Science Report 1: Prompt Engineering is Complicated and Contingent
- 不同评测标准显著影响大模型表现,选标准需对齐使用目标。
- 礼貌提示有时提升性能,有时反而降低,效果不确定。
- 提示策略需具体场景验证,通用公式不可靠,适合决策者参考。
本报告是系列简报的第一篇,旨在通过严谨测试帮助商业、教育和政策领导者理解与人工智能协作的技术细节。我们发现:第一,没有统一的衡量大语言模型(LLM)是否通过基准测试的标准,选择不同标准会显著影响其表现,标准应根据具体应用目标确定。第二,无法预先判断特定提示方法是否有助于回答某个问题——有时对模型保持礼貌能提升表现,有时则会降低;在某些情况下限制回答范围有助于性能提升,但在其他情况下反而有害。这表明,评估AI表现并非一刀切,特定提示策略(如礼貌性提问)也并非普遍有效。
原文摘要 · Abstract (English)
This is the first of a series of short reports that seek to help business, education, and policy leaders understand the technical details of working with AI through rigorous testing. In this report, we demonstrate two things: - There is no single standard for measuring whether a Large Language Model (LLM) passes a benchmark, and that choosing a standard has a big impact on how well the LLM does on that benchmark. The standard you choose will depend on your goals for using an LLM in a particular case. - It is hard to know in advance whether a particular prompting approach will help or harm the LLM's ability to answer any particular question. Specifically, we find that sometimes being polite to the LLM helps performance, and sometimes it lowers performance. We also find that constraining the AI's answers helps performance in some cases, though it may lower performance in other cases. Taken together, this suggests that benchmarking AI performance is not one-size-fits-all, and also that particular prompting formulas or approaches, like being polite to the AI, are not universally valuable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。