arXiv:2411.03395cs.HCcs.CL2024-11被引 8

大模型在乳腺癌专科诊疗中表现接近住院医生,但不及资深专家。

Exploring Large Language Models for Specialist-level Oncology Care

  • 用临床病例测试未微调的大模型,评估其乳腺癌诊疗能力。
  • 引入网络搜索和多阶段自检后,模型表现优于实习医生和主治医生。
  • 适合辅助医生决策,但尚未达到专家水平,仍需改进。

大型语言模型在编码临床知识和回应复杂医疗问题方面进展显著,但在亚专科或复杂医疗场景中的应用仍待探索。本文评估了未经特定领域微调的对话式诊断AI系统AMIE在乳腺肿瘤学领域的表现。我们构建了50个合成乳腺癌病例,涵盖治疗初治与耐药情况,模拟多学科肿瘤委员会的决策信息。开发了包含病例总结质量、治疗方案安全性及化疗、放疗、手术、激素治疗推荐等维度的临床评分标准。为提升性能,引入推理时网络检索获取最新临床知识,并通过多阶段自检流程优化回答。在自动评估和专科医师评估下,对比了AMIE与内科住院医师、肿瘤专科医师和普通肿瘤主治医师的表现。结果显示,AMIE优于住院医师和专科医师,但整体仍逊于资深主治医师,表明该系统在重要且具有挑战性的领域具备潜力,但仍需进一步研究方可用于实际临床决策支持。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown remarkable progress in encoding clinical knowledge and responding to complex medical queries with appropriate clinical reasoning. However, their applicability in subspecialist or complex medical settings remains underexplored. In this work, we probe the performance of AMIE, a research conversational diagnostic AI system, in the subspecialist domain of breast oncology care without specific fine-tuning to this challenging domain. To perform this evaluation, we curated a set of 50 synthetic breast cancer vignettes representing a range of treatment-naive and treatment-refractory cases and mirroring the key information available to a multidisciplinary tumor board for decision-making (openly released with this work). We developed a detailed clinical rubric for evaluating management plans, including axes such as the quality of case summarization, safety of the proposed care plan, and recommendations for chemotherapy, radiotherapy, surgery and hormonal therapy. To improve performance, we enhanced AMIE with the inference-time ability to perform web search retrieval to gather relevant and up-to-date clinical knowledge and refine its responses with a multi-stage self-critique pipeline. We compare response quality of AMIE with internal medicine trainees, oncology fellows, and general oncology attendings under both automated and specialist clinician evaluations. In our evaluations, AMIE outperformed trainees and fellows demonstrating the potential of the system in this challenging and important domain. We further demonstrate through qualitative examples, how systems such as AMIE might facilitate conversational interactions to assist clinicians in their decision making. However, AMIE's performance was overall inferior to attending oncologists suggesting that further research is needed prior to consideration of prospective uses.

大模型乳腺癌临床决策AI辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。