arXiv:2512.10995q-bio.QMcs.LG2025-12

用真实世界数据预测化疗方案失败,提升医生决策效率。

Boosted Random Forests for Predicting Treatment Failure of Chemotherapy Regimens

  • 基于临床记录构建特征向量,采用增强随机森林模型
  • 在五类常见癌症上达到80%准确率、75%F1分数
  • 兼顾性能与可解释性,适合临床医生直接使用

癌症患者常需经历漫长且痛苦的化疗疗程,包含多个连续治疗方案。治疗无效或不良反应可能导致方案中断或提前更换,给患者及其家庭带来显著身心与经济负担。本文基于肿瘤科电子病历系统中的真实世界证据(RWE),构建治疗失败预测模型。通过特征工程管道,从临床笔记、诊断和药物信息中提取独特特征向量,并在五种治疗失败率最高的主要癌症类型上进行建模。经过性能、复杂度与可解释性三维度设计探索,最终选定增强随机森林模型,在保持80%准确率和75% F1分数的同时降低模型复杂度,使结果更易被肿瘤医生理解和应用。

原文摘要 · Abstract (English)

Cancer patients may undergo lengthy and painful chemotherapy treatments, comprising several successive regimens or plans. Treatment inefficacy and other adverse events can lead to discontinuation (or failure) of these plans, or prematurely changing them, which results in a significant amount of physical, financial, and emotional toxicity to the patients and their families. In this work, we build treatment failure models based on the Real World Evidence (RWE) gathered from patients' profiles available in our oncology EMR/EHR system. We also describe our feature engineering pipeline, experimental methods, and valuable insights obtained about treatment failures from trained models. We report our findings on five primary cancer types with the most frequent treatment failures (or discontinuations) to build unique and novel feature vectors from the clinical notes, diagnoses, and medications that are available in our oncology EMR. After following a novel three axes - performance, complexity and explainability - design exploration framework, boosted random forests are selected because they provide a baseline accuracy of 80% and an F1 score of 75%, with reduced model complexity, thus making them more interpretable to and usable by oncologists.

医疗预测随机森林化疗可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。