arXiv:2601.07385cs.LG2026-01被引 1

用病历文本矩阵计算患者相似性,助力精准医疗

Computing patient similarity based on unstructured clinical notes

  • 将每位患者的病历文本聚合为嵌入矩阵,提取低秩特征
  • 在4267例乳腺癌患者数据上验证,支持个性化治疗推荐
  • 适用于临床史、治疗方案、不良反应等多维度相似性分析

临床病历包含诊断、治疗和预后等丰富但非结构化的信息,对精准医疗至关重要,却难以规模化利用。本文提出一种方法,将每位患者表示为由所有病历文本嵌入聚合而成的矩阵,基于其潜在低秩表示实现稳健的患者相似性计算。基于4,267名捷克乳腺癌患者的病历数据及马萨里克纪念癌症研究所提供的专家相似性标签,评估了多种基于矩阵的相似性度量,并分析其在临床史、治疗方案和不良事件等不同相似性维度上的优劣。结果表明,该方法在个性化治疗推荐或毒性预警等下游任务中具有实用价值。

原文摘要 · Abstract (English)

Clinical notes hold rich yet unstructured details about diagnoses, treatments, and outcomes that are vital to precision medicine but hard to exploit at scale. We introduce a method that represents each patient as a matrix built from aggregated embeddings of all their notes, enabling robust patient similarity computation based on their latent low-rank representations. Using clinical notes of 4,267 Czech breast-cancer patients and expert similarity labels from Masaryk Memorial Cancer Institute, we evaluate several matrix-based similarity measures and analyze their strengths and limitations across different similarity facets, such as clinical history, treatment, and adverse events. The results demonstrate the usefulness of the presented method for downstream tasks, such as personalized therapy recommendations or toxicity warnings.

患者相似性临床文本精准医疗低秩表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。