小改动提示词对复杂预测任务提升有限,反而可能让模型更不准。
Prompt Engineering Large Language Models' Forecasting Capabilities
- 测试38种提示词,仅基础率引用带来轻微提升
- 鼓励贝叶斯推理的提示反而显著降低预测准确率
- 简单提示工程难以支撑复杂预测,需更专门方法
大语言模型性能可通过多种方式提升,但微调或高级工具使用成本高。提示工程成本低且适用于简单任务,但其在复杂领域如预测中的有效性尚不明确。本文在Claude 3.5 Sonnet、Claude 3.5 Haiku、GPT-4o和Llama 3.1 405B上测试了38种提示词;第二阶段引入复合提示及外部来源提示,并纳入o1和o1-mini推理模型。结果表明,多数提示仅带来可忽略的改进,仅有基础率参考略有效果。令人意外的是,鼓励模型进行贝叶斯推理的策略反而显著降低准确性。这表明,在复杂任务如预测中,基础提示优化作用有限,可能需要更稳健或专用的技术才能实现显著性能提升。
原文摘要 · Abstract (English)
Large language model performance can be improved in a large number of ways. Many such techniques, like fine-tuning or advanced tool usage, are time-intensive and expensive. Although prompt engineering is significantly cheaper and often works for simpler tasks, it remains unclear whether prompt engineering suffices for more complex domains like forecasting. Here we show that small prompt modifications rarely boost forecasting accuracy beyond a minimal baseline. In our first study, we tested 38 prompts across Claude 3.5 Sonnet, Claude 3.5 Haiku, GPT-4o, and Llama 3.1 405B. In our second, we introduced compound prompts and prompts from external sources, also including the reasoning models o1 and o1-mini. Our results show that most prompts lead to negligible gains, although references to base rates yield slight benefits. Surprisingly, some strategies showed strong negative effects on accuracy: especially encouraging the model to engage in Bayesian reasoning. These results suggest that, in the context of complex tasks like forecasting, basic prompt refinements alone offer limited gains, implying that more robust or specialized techniques may be required for substantial performance improvements in AI forecasting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。