微调开源大模型提升放疗任务表现,验证其临床应用潜力
Fine-Tuning Open-Source Large Language Models to Improve Their Performance on Radiation Oncology Tasks: A Feasibility Study to Investigate Their Potential Clinical Applications in Radiation Oncology
- 用放疗病例数据微调LLaMA2和Mistral模型,构建三类任务对
- 微调后模型在三项任务中准确率显著提升,临床可接受率超60%
- 适合医疗AI研究者与放疗领域从业者关注放疗自动化方案
放射肿瘤学临床实践依赖大量文本信息的动态交互。大语言模型虽在处理复杂文本方面表现突出,但在放疗等特定领域的应用仍较少。本研究旨在探究基于领域知识微调大模型能否提升其在三个任务上的表现:治疗方案生成、治疗方式选择(光子、质子、电子或近距离放疗)以及ICD-10编码预测。从15,724例患者中筛选出7,903例具有单一诊断记录且明确主治疗方案的病例,经预处理与人工标注后构建训练对。使用开源的LLaMA2-7B与Mistral-7B模型,结合低秩近似方法进行微调。报告了微调前后模型的准确率与ROUGE-1得分,并由放疗医师对治疗方案生成任务进行临床评估,其余任务采用精确率、召回率与F1分数评价。采用单侧威尔科克斯符号秩检验分析结果。结果显示,微调模型在所有任务中均显著优于原始模型(p ≤ 0.001)。临床评估表明,超过60%的微调模型生成的治疗方案具备临床可接受性。各任务的精确率、召回率与F1分数均有提升。
原文摘要 · Abstract (English)
Background: The radiation oncology clinical practice involves many steps relying on the dynamic interplay of abundant text data. Large language models have displayed remarkable capabilities in processing complex text information. But their direct applications in specific fields like radiation oncology remain underexplored. Purpose: This study aims to investigate whether fine-tuning LLMs with domain knowledge can improve the performance on Task (1) treatment regimen generation, Task (2) treatment modality selection (photon, proton, electron, or brachytherapy), and Task (3) ICD-10 code prediction in radiation oncology. Methods: Data for 15,724 patient cases were extracted. Cases where patients had a single diagnostic record, and a clearly identifiable primary treatment plan were selected for preprocessing and manual annotation to have 7,903 cases of the patient diagnosis, treatment plan, treatment modality, and ICD-10 code. Each case was used to construct a pair consisting of patient diagnostics details and an answer (treatment regimen, treatment modality, or ICD-10 code respectively) for the supervised fine-tuning of these three tasks. Open source LLaMA2-7B and Mistral-7B models were utilized for the fine-tuning with the Low-Rank Approximations method. Accuracy and ROUGE-1 score were reported for the fine-tuned models and original models. Clinical evaluation was performed on Task (1) by radiation oncologists, while precision, recall, and F-1 score were evaluated for Task (2) and (3). One-sided Wilcoxon signed-rank tests were used to statistically analyze the results. Results: Fine-tuned LLMs outperformed original LLMs across all tasks with p-value <= 0.001. Clinical evaluation demonstrated that over 60% of the fine-tuned LLMs-generated treatment regimens were clinically acceptable. Precision, recall, and F1-score showed improved performance of fine-tuned LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。