arXiv:2606.02892cs.LG2026-06

融合病历、病理和医生笔记,提升乳腺癌复发预测准确率

Multi-Modal Machine Learning for Breast Cancer Recurrence Prediction

论文配图:Multi-Modal Machine Learning for Breast Cancer Recurrence Prediction
图 1 · 摘自论文原文
  • 从自由文本中提取肿瘤特征,补全结构化数据
  • 多模态输入比单一数据源预测准确率更高
  • 适合临床风险评估与个性化随访决策研究者

乳腺癌复发是幸存者长期死亡的主要原因,及时准确的风险评估对随访和治疗规划至关重要。传统预测模型通常仅依赖结构化或非结构化数据之一,难以捕捉完整的临床信息。本研究探讨整合治疗记录、病理报告和医生笔记等多模态临床数据对复发预测的影响。通过规则匹配的正则表达式提取机制与基于优先级的冲突协调策略,有效从自由文本病理描述中恢复关键肿瘤特征,补充结构化记录。同时,与以往乳腺癌研究常用的特征集进行对比,评估多模态融合的价值。在多种机器学习模型上比较单源与多模态输入的表现。结果显示,多模态集成在各类模型中均显著提升预测准确性。

原文摘要 · Abstract (English)

Breast cancer recurrence, a leading cause of long-term mortality among survivors, requires timely and accurate risk assessment to guide follow-up care and treatment planning. Traditional predictive models, often limited to either structured or unstructured data alone, struggle to capture the full clinical context. This study examines the impact of integrating multi-modal clinical data, including treatment records, pathology reports, and clinician notes, on recurrence prediction. By integrating a rule-based regular expression extraction mechanism with a rigorous precedence-based conflict reconciliation strategy, our approach effectively recovers definitive tumor characteristics from free-text pathology narratives to augment structured records. We also benchmark performance against commonly used feature sets from prior breast cancer studies to assess the added value of multi-modal integration. Single-source and multi-modal inputs are evaluated across a range of machine learning models. Results show that multi-modal integration consistently improves predictive accuracy compared to single-modal methods.

乳腺癌多模态复发预测临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。