arXiv:2412.10622physics.med-phcs.AI2024-12被引 11

大模型在放疗物理题上表现超人,但需提示才能发挥推理能力。

A recent evaluation on the performance of LLMs on radiation oncology physics using questions of randomly shuffled options

  • 用打乱选项的试题测试大模型,检验其推理能力。
  • o1-preview模型答题正确率超医生集体判断,其他模型性能也达专家水平。
  • 先解释再推理的提示策略能显著提升部分模型的解题能力。

目的:评估最新发布的大语言模型(LLMs)在放疗物理领域的问题解答能力。方法:使用100道由资深物理学家设计的多选题,对答案选项进行随机打乱生成新题集。测试了五款在2024年9月30日前发布的模型:OpenAI o1-preview、GPT-4o、LLaMA 3.1 (405B)、Gemini 1.5 Pro 和 Claude 3.5 Sonnet。为评估其演绎推理能力,将正确选项替换为“以上皆非”,并采用“先解释”和“分步推理”提示策略,观察是否改善表现。结果:所有模型在原题中均表现出专家级水平,其中o1-preview在多数情况下超越医疗物理师集体判断;当正确选项被替换为“以上皆非”后,各模型表现显著下降,表明仍有改进空间。而“先解释”和“分步推理”提示显著提升了LLaMA 3.1 (405B)、Gemini 1.5 Pro 和 Claude 3.5 Sonnet 的推理表现。结论:近期发布的这些大模型在放疗物理问题上已具备专家级能力,展现出在放疗物理教育与培训中的巨大应用潜力。

原文摘要 · Abstract (English)

Purpose: We present an updated study evaluating the performance of large language models (LLMs) in answering radiation oncology physics questions, focusing on the recently released models. Methods: A set of 100 multiple-choice radiation oncology physics questions, previously created by a well-experienced physicist, was used for this study. The answer options of the questions were randomly shuffled to create "new" exam sets. Five LLMs -- OpenAI o1-preview, GPT-4o, LLaMA 3.1 (405B), Gemini 1.5 Pro, and Claude 3.5 Sonnet -- with the versions released before September 30, 2024, were queried using these new exam sets. To evaluate their deductive reasoning ability, the correct answer options in the questions were replaced with "None of the above." Then, the explain-first and step-by-step instruction prompts were used to test if this strategy improved their reasoning ability. The performance of the LLMs was compared with the answers from medical physicists. Results: All models demonstrated expert-level performance on these questions, with o1-preview even surpassing medical physicists with a majority vote. When replacing the correct answer options with 'None of the above', all models exhibited a considerable decline in performance, suggesting room for improvement. The explain-first and step-by-step instruction prompts helped enhance the reasoning ability of the LLaMA 3.1 (405B), Gemini 1.5 Pro, and Claude 3.5 Sonnet models. Conclusion: These recently released LLMs demonstrated expert-level performance in answering radiation oncology physics questions, exhibiting great potential to assist in radiation oncology physics education and training.

大模型放疗物理推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。