o1模型在医学任务中表现优于GPT-4,但存在幻觉和评估不一致问题。
A Preliminary Study of o1 in Medicine: Are We Closer to an AI Doctor?
- 采用强化学习实现内部思维链,提升医学推理能力
- 在19个数据集上平均准确率超GPT-4 6.2%,复杂题达6.6%
- 适合临床研究者关注其推理优势与评估缺陷
大型语言模型(LLMs)在多个领域展现出强大能力,推动认知学习边界。最新模型OpenAI o1首次通过强化学习策略实现内化思维链,显著提升通用语言任务表现。然而其在医学等专业领域的效能尚不明确。本报告系统评估o1在医学场景中的理解、推理与多语言能力,涵盖6项任务及37个医学数据集,包括基于《新英格兰医学杂志》和《柳叶刀》的两个新构建高难度问答任务,相比传统基准(如MedQA)更具临床相关性。结果显示,o1在19个数据集上平均准确率比GPT-4高出6.2%,在新设复杂问答任务上提升6.6%。分析表明,增强推理能力显著提升医学指令理解和复杂病例推断能力。但同时发现模型存在幻觉、多语言表现不稳及评估指标不一致等问题。原始数据与模型输出已公开于https://ucsc-vlaa.github.io/o1_medicine/,供后续研究使用。
原文摘要 · Abstract (English)
Large language models (LLMs) have exhibited remarkable capabilities across various domains and tasks, pushing the boundaries of our knowledge in learning and cognition. The latest model, OpenAI's o1, stands out as the first LLM with an internalized chain-of-thought technique using reinforcement learning strategies. While it has demonstrated surprisingly strong capabilities on various general language tasks, its performance in specialized fields such as medicine remains unknown. To this end, this report provides a comprehensive exploration of o1 on different medical scenarios, examining 3 key aspects: understanding, reasoning, and multilinguality. Specifically, our evaluation encompasses 6 tasks using data from 37 medical datasets, including two newly constructed and more challenging question-answering (QA) tasks based on professional medical quizzes from the New England Journal of Medicine (NEJM) and The Lancet. These datasets offer greater clinical relevance compared to standard medical QA benchmarks such as MedQA, translating more effectively into real-world clinical utility. Our analysis of o1 suggests that the enhanced reasoning ability of LLMs may (significantly) benefit their capability to understand various medical instructions and reason through complex clinical scenarios. Notably, o1 surpasses the previous GPT-4 in accuracy by an average of 6.2% and 6.6% across 19 datasets and two newly created complex QA scenarios. But meanwhile, we identify several weaknesses in both the model capability and the existing evaluation protocols, including hallucination, inconsistent multilingual ability, and discrepant metrics for evaluation. We release our raw data and model outputs at https://ucsc-vlaa.github.io/o1_medicine/ for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。