用大模型从临床笔记中提取乳腺癌表型,效果媲美传统知识库方法。
Extracting Breast Cancer Phenotypes from Clinical Notes: Comparing LLMs with Classical Ontology Methods

- 基于大模型构建信息抽取框架,自动解析医生自然语言记录
- 提取乳腺癌相关表型准确率与经典本体方法相当
- 训练后可轻松适配其他癌症类型,扩展性强
肿瘤学电子病历(EMR)中大量数据以未结构化的医生笔记形式存在,包括化疗结果、生物标志物、肿瘤位置、大小及生长模式等关键信息。临床研究表明,多数肿瘤科医生更倾向于用自然语言描述这些重要信息,而非填写结构化字段。本文提出一种基于大语言模型(LLM)的框架,用于从临床笔记中提取乳腺癌相关表型,并与基于NCIt本体注释器的经典知识驱动方法进行对比。实验结果表明,该LLM框架在提取表型方面表现良好,准确率与传统本体方法相当;更重要的是,一旦训练完成,可快速微调以适用于其他癌症类型和疾病,具备良好的泛化能力。
原文摘要 · Abstract (English)
A significant amount of data held in Oncology Electronic Medical Records (EMRs) is contained in unstructured provider notes -- including but not limited to the chemotherapy (or cancer treatment) outcome, different biomarkers, the tumor's location, sizes, and growth patterns of a patient. The clinical studies show that the majority of oncologists are comfortable providing these valuable insights in their notes in a natural language rather than the relevant structured fields of an EMR. The major contribution of this research is to report an LLM-based framework to process provider notes and extract valuable medical knowledge and phenotype mentioned above, with a focus on the domain of oncology. In this paper, we focus on extracting phenotypes related to breast cancer using our LLM framework, and then compare its performance with earlier works that used knowledge-driven annotation system, paired with the NCIt Ontology Annotator. The results of the study show that an LLM-based information extraction framework can be easily adapted to extract phenotypes with an accuracy that is comparable to the classical ontology-based methods. However, once trained, they could be easily fine-tuned to cater for other cancer types and diseases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。