测试大模型能否准确识别乳腺癌放疗副作用,发现其对细节敏感且易漏罕见长期副作用。
Can Language Models Identify Side Effects of Breast Cancer Radiation Treatments?
- 构建对比临床场景,用7个微调模型测试不同提示下的副作用生成能力。
- 多数模型在稀有和长期副作用上召回率低,且对放疗方案微小差异敏感。
- 结合医生整理的参考清单可显著提升结果可靠性,适合临床辅助应用。
准确向癌症幸存者传达治疗副作用至关重要,尤其在知情同意等场景中,但因临床知识不足及电子病历系统碎片化,这一任务仍具挑战。大型语言模型(LLMs)可能在此方面提供帮助,但其在肿瘤康复场景中的可靠性尚不明确。本文提出一个面向部署的应力测试框架,评估乳腺癌放疗相关副作用的LLM生成表现。基于21名乳腺癌患者资料,构建仅在放疗方案上不同的配对临床情景,测试7个指令微调的LLM在多种提示策略下的表现。将输出与由超过七位乳腺放射肿瘤学家团队制定的两所顶级医学中心知情同意文件参考标准进行比较,该标准涵盖放疗剂量-分次、照射野及部位对应的毒性反应,按频率和发生时间划分。结果显示,模型对文档微小变化敏感,存在精确率与召回率权衡,且对罕见及长期副作用系统性遗漏。单独使用时,限制副作用数量会降低精确率;而将输出锚定于医生整理的参考清单,能显著提升可靠性与鲁棒性。研究揭示了LLM在肿瘤学应用中的重要局限,并为更安全、更信息丰富的康复导向应用提供了实用设计建议。
原文摘要 · Abstract (English)
Accurately communicating the side effects of cancer treatments to cancer survivors is critical, particularly in settings such as informed consent, where clinicians must clearly and comprehensively convey potential treatment toxicities. However, this task remains challenging due to clinical knowledge deficits about adverse treatment effects and fragmentation across electronic health record (EHR) systems. Large language models (LLMs) have the potential to assist in this task, though their reliability in oncology survivorship contexts remains poorly understood. We present a deployment-oriented stress-testing framework for evaluating LLM-generated radiation side effect lists in breast cancer treatment and survivorship care. Using 21 breast cancer patient profiles, we construct paired patient clinical scenarios that differ only in radiotherapy regimens to evaluate seven instruction-tuned LLMs under multiple prompting regimes. We then compare LLM outputs to a clinician-curated reference derived from informed consent documents at two major academic medical centers and developed by a team including more than seven breast radiation oncologists. The reference maps radiation dose-fractionation, fields, and locations to associated toxicities, broken down by frequency and temporal onset. Across models, we reveal sensitivity to minor documentation changes, trade-offs between precision and recall, and systematic under-recall of rare and long-term side effects. When used alone, constraints on the number of side effects generated reduce precision, and grounding outputs in clinician-curated side effect lists substantially improves reliability and robustness. These findings highlight important limitations of LLM use in oncology and suggest practical design choices for safer and more informative survivorship-focused applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。