o1-preview在医疗任务中表现超越GPT-4,但需权衡成本与效果。
From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond
- o1-preview通过内置推理机制,无需外部提示即可达到顶尖性能。
- 少样本提示会削弱o1表现,说明其不依赖上下文学习。
- 适合追求高精度且可接受高成本的医疗智能应用开发者。
运行时引导策略如Medprompt对大型语言模型(LLMs)在挑战性任务中实现卓越表现具有价值。Medprompt表明,通过提示激发链式思维和集成推理,通用LLM可在医学等专业领域达到领先水平。OpenAI的o1-preview代表新范式,模型在生成最终回答前进行运行时推理。我们系统评估o1-preview在多种医疗基准上的表现。结果表明,即使无提示技术,o1-preview仍显著优于GPT-4系列搭配Medprompt的表现。我们进一步研究经典提示工程(如Medprompt)在推理原生模型中的有效性,发现少样本提示会损害o1性能,暗示上下文学习对这类模型不再有效。尽管集成仍可行,但资源消耗大,需精细优化成本-性能平衡。跨策略的成本与准确率分析揭示了帕累托前沿:GPT-4o成本更低,而o1-preview在高成本下达成最先进性能。尽管o1-preview表现顶尖,但搭配Medprompt策略的GPT-4o在特定场景仍具价值。此外,o1-preview在多数现有医疗基准上已接近饱和,凸显构建更难基准的必要性。最后,论文反思了大模型推理时计算的未来方向。
原文摘要 · Abstract (English)
Run-time steering strategies like Medprompt are valuable for guiding large language models (LLMs) to top performance on challenging tasks. Medprompt demonstrates that a general LLM can be focused to deliver state-of-the-art performance on specialized domains like medicine by using a prompt to elicit a run-time strategy involving chain of thought reasoning and ensembling. OpenAI's o1-preview model represents a new paradigm, where a model is designed to do run-time reasoning before generating final responses. We seek to understand the behavior of o1-preview on a diverse set of medical challenge problem benchmarks. Following on the Medprompt study with GPT-4, we systematically evaluate the o1-preview model across various medical benchmarks. Notably, even without prompting techniques, o1-preview largely outperforms the GPT-4 series with Medprompt. We further systematically study the efficacy of classic prompt engineering strategies, as represented by Medprompt, within the new paradigm of reasoning models. We found that few-shot prompting hinders o1's performance, suggesting that in-context learning may no longer be an effective steering approach for reasoning-native models. While ensembling remains viable, it is resource-intensive and requires careful cost-performance optimization. Our cost and accuracy analysis across run-time strategies reveals a Pareto frontier, with GPT-4o representing a more affordable option and o1-preview achieving state-of-the-art performance at higher cost. Although o1-preview offers top performance, GPT-4o with steering strategies like Medprompt retains value in specific contexts. Moreover, we note that the o1-preview model has reached near-saturation on many existing medical benchmarks, underscoring the need for new, challenging benchmarks. We close with reflections on general directions for inference-time computation with LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。