arXiv:2501.02346physics.med-phcs.AI2025-01被引 12

GPT-4在放疗医学领域表现优异,助力临床决策支持。

Exploring the Capabilities and Limitations of Large Language Models for Radiation Oncology Decision Support

  • 通过专业试题测试,评估GPT-4在放疗物理领域的表现。
  • 在临床放疗考试中准确率达74.57%,结构命名重标准确超96%。
  • 适合医疗AI研究者与放疗临床医生参考其应用边界。

随着大语言模型(LLM)在决策支持系统中的快速集成,大型系统正经历显著变革。如同其他医学领域,放疗学界对GPT-4等大模型的兴趣日益增长。一项针对放疗物理这一高度专业领域的100题专项测试显示,GPT-4在性能上优于其他大模型。在更广泛的临床放疗领域,通过美国放射学会(ACR)放疗住院医师培训(TXIT)考试对GPT-4进行基准测试,其准确率达到74.57%。此外,依据AAPM TG-263报告对结构名称进行重新标注的任务中,其准确率超过96%。这些研究揭示了大语言模型在放疗决策支持中的潜力。然而,随着对大模型在通用医疗应用中潜力与局限性关注持续上升,其在放疗领域的具体能力与限制尚未得到充分探索。

原文摘要 · Abstract (English)

Thanks to the rapidly evolving integration of LLMs into decision-support tools, a significant transformation is happening across large-scale systems. Like other medical fields, the use of LLMs such as GPT-4 is gaining increasing interest in radiation oncology as well. An attempt to assess GPT-4's performance in radiation oncology was made via a dedicated 100-question examination on the highly specialized topic of radiation oncology physics, revealing GPT-4's superiority over other LLMs. GPT-4's performance on a broader field of clinical radiation oncology is further benchmarked by the ACR Radiation Oncology In-Training (TXIT) exam where GPT-4 achieved a high accuracy of 74.57%. Its performance on re-labelling structure names in accordance with the AAPM TG-263 report has also been benchmarked, achieving above 96% accuracies. Such studies shed light on the potential of LLMs in radiation oncology. As interest in the potential and constraints of LLMs in general healthcare applications continues to rise5, the capabilities and limitations of LLMs in radiation oncology decision support have not yet been fully explored.

放疗决策大模型GPT-4医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。