大模型仍需提示优化,用大模型优化提示更有效。
Revisiting Prompt Optimization with Large Reasoning Models-A Case Study on Event Extraction
- 用大模型做提示优化,比直接使用更精准
- 事件抽取任务中,大模型仍需提示调优才能提效
- 适合需要高精度指令设计的复杂任务研究者
大型推理模型(LRMs)如 DeepSeek-R1 和 OpenAI o1 在各类推理任务中展现出卓越能力,其生成和推理中间思考过程的能力,引发学界讨论:是否已无需大量提示工程即可准确理解指令并输出结果。本文以事件抽取这一结构化任务为案例,系统研究该问题。实验对比了两种 LRMs(DeepSeek-R1、o1)与两种通用大语言模型(GPT-4o、GPT-4.5)在作为任务模型或提示优化器时的表现。结果显示,在事件抽取这类复杂任务中,即使使用 LRM 作为任务模型,仍能从提示优化中获益;而使用 LRM 作为提示优化器,可生成更有效的提示。该结论也适用于其他任务。最后,我们对 LRM 常见错误进行分析,指出其在细化任务指令和事件规范方面具备稳定性和一致性。
原文摘要 · Abstract (English)
Large Reasoning Models (LRMs) such as DeepSeek-R1 and OpenAI o1 have demonstrated remarkable capabilities in various reasoning tasks. Their strong capability to generate and reason over intermediate thoughts has also led to arguments that they may no longer require extensive prompt engineering or optimization to interpret human instructions and produce accurate outputs. In this work, we aim to systematically study this open question, using the structured task of event extraction for a case study. We experimented with two LRMs (DeepSeek-R1 and o1) and two general-purpose Large Language Models (LLMs) (GPT-4o and GPT-4.5), when they were used as task models or prompt optimizers. Our results show that on tasks as complicated as event extraction, LRMs as task models still benefit from prompt optimization, and that using LRMs as prompt optimizers yields more effective prompts. Our finding also generalizes to tasks beyond event extraction. Finally, we provide an error analysis of common errors made by LRMs and highlight the stability and consistency of LRMs in refining task instructions and event guidelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。