arXiv:2410.19826cs.LG2024-10被引 1

用大模型将癌症病历转为标准数据,提升临床试验匹配效率

Novel Development of LLM Driven mCODE Data Model for Improved Clinical Trial Matching to Enable Standardization and Interoperability in Oncology Research

  • 用大模型自动提取病历文本,生成符合mCODE标准的患者数据档案
  • 模型在数千条数据上准确率达92%,远超GPT-4等现有模型的77%平均值
  • 适合医疗数据标准化、跨系统协作及精准肿瘤治疗研究者使用

每年因癌症诊疗数据缺乏标准化与互操作性,导致诊断延迟、成本飙升,2023年全国癌症支出已超2080亿美元。传统临床试验入组与诊疗方式依赖人工,耗时且缺乏数据驱动。本文提出一种新框架,利用先进大模型与计算机工程,推动癌症领域数据标准化、互操作性与跨系统共享。通过FHIR资源模型与大模型生成的mCODE档案,实现患者信息高效传递。方法将非结构化病历、PDF、自由文本和病程记录转化为丰富mCODE档案,支持与新型基于AI/ML的临床试验匹配引擎无缝集成。结果显示,训练后的模型在数千例患者数据上准确率峰值达92%;对SNOMED-CT、LOINC、RxNorm编码识别率分别为87%、90%、84%,显著优于当前主流模型如GPT-4与Claude 3.5的平均77%表现。该框架证明了其在提升癌症诊疗效率与个性化治疗方面的潜力。

原文摘要 · Abstract (English)

Each year, the lack of efficient data standardization and interoperability in cancer care contributes to the severe lack of timely and effective diagnosis, while constantly adding to the burden of cost, with cancer costs nationally reaching over $208 billion in 2023 alone. Traditional methods regarding clinical trial enrollment and clinical care in oncology are often manual, time-consuming, and lack a data-driven approach. This paper presents a novel framework to streamline standardization, interoperability, and exchange of cancer domains and enhance the integration of oncology-based EHRs across disparate healthcare systems. This paper utilizes advanced LLMs and Computer Engineering to streamline cancer clinical trials and discovery. By utilizing FHIR's resource-based approach and LLM-generated mCODE profiles, we ensure timely, accurate, and efficient sharing of patient information across disparate healthcare systems. Our methodology involves transforming unstructured patient treatment data, PDFs, free-text information, and progress notes into enriched mCODE profiles, facilitating seamless integration with our novel AI and ML-based clinical trial matching engine. The results of this study show a significant improvement in data standardization, with accuracy rates of our trained LLM peaking at over 92% with datasets consisting of thousands of patient data. Additionally, our LLM demonstrated an accuracy rate of 87% for SNOMED-CT, 90% for LOINC, and 84% for RxNorm codes. This trumps the current status quo, with LLMs such as GPT-4 and Claude's 3.5 peaking at an average of 77%. This paper successfully underscores the potential of our standardization and interoperability framework, paving the way for more efficient and personalized cancer treatment.

大模型医疗数据临床试验标准化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。