测试给AI打赏或威胁是否真能提升表现,结果发现效果不明显。
Prompting Science Report 3: I'll pay you or I'll kill you -- but will you care?
- 通过打赏或威胁提示,测试对AI性能影响
- 在GPQA和MMLU-Pro上均未显著提升表现
- 个别问题结果差异大,但无法提前预判
本报告是系列短报第三篇,旨在帮助商业、教育和政策领导者理解与AI协作的技术细节。我们检验了两种常见提示策略:一是承诺打赏模型,二是威胁模型。打赏是普遍推荐的提升方法,而谷歌创始人谢尔盖·布林曾表示‘模型在被威胁时表现更好’,本文对此进行实证检验。我们在GPQA(Rein et al. 2024)和MMLU-Pro(Wang et al. 2024)两个基准上评估模型表现,发现:威胁或打赏对整体性能无显著影响;但提示方式变化可在单题层面显著影响结果。然而,难以提前判断某种提示策略会对特定问题产生帮助还是损害。这表明简单提示技巧对复杂问题可能不如预期有效,但如前所述(Meincke et al. 2025a),提示策略对个别问题仍可带来显著差异。
原文摘要 · Abstract (English)
This is the third in a series of short reports that seek to help business, education, and policy leaders understand the technical details of working with AI through rigorous testing. In this report, we investigate two commonly held prompting beliefs: a) offering to tip the AI model and b) threatening the AI model. Tipping was a commonly shared tactic for improving AI performance and threats have been endorsed by Google Founder Sergey Brin (All-In, May 2025, 8:20) who observed that 'models tend to do better if you threaten them,' a claim we subject to empirical testing here. We evaluate model performance on GPQA (Rein et al. 2024) and MMLU-Pro (Wang et al. 2024). We demonstrate two things: - Threatening or tipping a model generally has no significant effect on benchmark performance. - Prompt variations can significantly affect performance on a per-question level. However, it is hard to know in advance whether a particular prompting approach will help or harm the LLM's ability to answer any particular question. Taken together, this suggests that simple prompting variations might not be as effective as previously assumed, especially for difficult problems. However, as reported previously (Meincke et al. 2025a), prompting approaches can yield significantly different results for individual questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。