用大模型从病历中提取关键信息,提前预测化疗效果
Leveraging Large Language Models and Survival Analysis for Early Prediction of Chemotherapy Outcomes
- 用大模型和本体技术从病历中自动提取表型和治疗结果标签
- 在乳腺癌数据上实现73% C-index,关键时间点预测准确率超70%
- 方法可推广至其他癌症类型,助力个性化化疗决策
化疗成本高且副作用严重,亟需早期预测治疗效果以优化管理。现有模型因缺乏明确表型和治疗结局标签(如肿瘤进展、毒性)面临挑战。本研究针对乳腺癌(高发且治疗反应差异大),利用大语言模型(LLMs)和本体技术从电子病历文本中提取表型与结局标签。数据包含生命体征、人口学、分期、生物标志物及功能评分;化疗方案基于NCCN指南与NIH标准提取并验证。采用随机生存森林建模,时间到失败的预测C-index达73%,在特定时间点作为分类器时准确率与F1值均超过70%。校准曲线验证了结果可靠性,并扩展至四种其他癌症类型。研究表明,基于大模型的临床数据提取可实现早期疗效预测,支持个体化治疗方案制定。
原文摘要 · Abstract (English)
Chemotherapy for cancer treatment is costly and accompanied by severe side effects, highlighting the critical need for early prediction of treatment outcomes to improve patient management and informed decision-making. Predictive models for chemotherapy outcomes using real-world data face challenges, including the absence of explicit phenotypes and treatment outcome labels such as cancer progression and toxicity. This study addresses these challenges by employing Large Language Models (LLMs) and ontology-based techniques for phenotypes and outcome label extraction from patient notes. We focused on one of the most frequently occurring cancers, breast cancer, due to its high prevalence and significant variability in patient response to treatment, making it a critical area for improving predictive modeling. The dataset included features such as vitals, demographics, staging, biomarkers, and performance scales. Drug regimens and their combinations were extracted from the chemotherapy plans in the EMR data and shortlisted based on NCCN guidelines, verified with NIH standards, and analyzed through survival modeling. The proposed approach significantly reduced phenotypes sparsity and improved predictive accuracy. Random Survival Forest was used to predict time-to-failure, achieving a C-index of 73%, and utilized as a classifier at a specific time point to predict treatment outcomes, with accuracy and F1 scores above 70%. The outcome probabilities were validated for reliability by calibration curves. We extended our approach to four other cancer types. This research highlights the potential of early prediction of treatment outcomes using LLM-based clinical data extraction enabling personalized treatment plans with better patient outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。