arXiv:2508.21777cs.CVcs.AI2025-08被引 5

GPT-5在放疗科考试中准确率达92.8%,但复杂病例仍需专家把关。

Benchmarking GPT-5 in Radiation Oncology: Measurable Gains, but Persistent Need for Expert Oversight

  • 用真实临床案例和考试题评估GPT-5放疗决策能力
  • 多选题正确率92.8%,治疗方案正确性评分3.24/4
  • 虽幻觉罕见,但复杂病例仍有实质性错误,需专家审核

大型语言模型(LLM)在临床决策支持中展现巨大潜力。GPT-5是专为肿瘤学设计的新一代模型。通过两项互补基准测试评估性能:(i) 美国放射学会放疗住院医师考试(TXIT, 2021),含300道选择题;(ii) 60个真实放疗病例,涵盖多种瘤种与治疗指征。对病例评估中,GPT-5被要求生成简洁治疗方案,由四位认证放疗科医生评分正确性、完整性及幻觉情况。组间一致性以Fleiss' kappa衡量。结果:在TXIT测试中,GPT-5平均准确率为92.8%,优于GPT-4(78.8%)和GPT-3.5(62.1%),尤其在剂量与诊断领域提升显著。在病例评估中,治疗建议正确性均值为3.24/4(95% CI: 3.11–3.38),完整性为3.59/4(95% CI: 3.49–3.69)。幻觉极少,无一例达多数共识。组间一致性低(Fleiss' kappa=0.083),反映临床判断固有差异。错误集中于需精确试验知识或详细临床适配的复杂场景。讨论:尽管GPT-5在放疗多选题中明显优于前代模型,其生成的真实治疗建议虽整体表现良好,但仍存在改进空间。尽管幻觉少见,实质性错误仍提示其推荐必须经严格专家审核后方可用于临床。

原文摘要 · Abstract (English)

Introduction: Large language models (LLM) have shown great potential in clinical decision support. GPT-5 is a novel LLM system that has been specifically marketed towards oncology use. Methods: Performance was assessed using two complementary benchmarks: (i) the ACR Radiation Oncology In-Training Examination (TXIT, 2021), comprising 300 multiple-choice items, and (ii) a curated set of 60 authentic radiation oncologic vignettes representing diverse disease sites and treatment indications. For the vignette evaluation, GPT-5 was instructed to generate concise therapeutic plans. Four board-certified radiation oncologists rated correctness, comprehensiveness, and hallucinations. Inter-rater reliability was quantified using Fleiss' \k{appa}. Results: On the TXIT benchmark, GPT-5 achieved a mean accuracy of 92.8%, outperforming GPT-4 (78.8%) and GPT-3.5 (62.1%). Domain-specific gains were most pronounced in Dose and Diagnosis. In the vignette evaluation, GPT-5's treatment recommendations were rated highly for correctness (mean 3.24/4, 95% CI: 3.11-3.38) and comprehensiveness (3.59/4, 95% CI: 3.49-3.69). Hallucinations were rare with no case reaching majority consensus for their presence. Inter-rater agreement was low (Fleiss' \k{appa} 0.083 for correctness), reflecting inherent variability in clinical judgment. Errors clustered in complex scenarios requiring precise trial knowledge or detailed clinical adaptation. Discussion: GPT-5 clearly outperformed prior model variants on the radiation oncology multiple-choice benchmark. Although GPT-5 exhibited favorable performance in generating real-world radiation oncology treatment recommendations, correctness ratings indicate room for further improvement. While hallucinations were infrequent, the presence of substantive errors underscores that GPT-5-generated recommendations require rigorous expert oversight before clinical implementation.

大模型放疗临床决策AI医疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。