arXiv:2502.15996cs.CLcs.AI2025-02被引 1

基于临床文本的智能信息提取模型,提升疾病预测与患者分组效果

Med-gte-hybrid: A contextual embedding transformer model for extracting actionable information from clinical texts

  • 融合对比学习与去噪自编码器,优化医学语境嵌入
  • 在MIMIC-IV数据上实现肾病预后、肌酐清除率等任务的精准预测
  • 适用于多种医疗场景,助力个性化诊疗决策

我们提出一种新型上下文嵌入模型 med-gte-hybrid,源自 gte-large 句子变换器,用于从非结构化临床文本中提取可操作信息。该模型通过结合对比学习与去噪自编码器进行调优。为评估性能,我们在 MIMIC-IV 数据集中大样本患者队列上测试了多项临床预测任务,包括慢性肾病(CKD)患者预后、估算肾小球滤过率(eGFR)预测及患者死亡率预测。此外,我们证明 med-gte-hybrid 在患者分层、聚类和文本检索方面表现优异,在 Massive Text Embedding Benchmark(MTEB)上超越现有最优模型。尽管部分评估聚焦于 CKD,但该混合调优策略可迁移至其他医学领域,具备提升临床决策与个性化治疗路径的潜力。

原文摘要 · Abstract (English)

We introduce a novel contextual embedding model med-gte-hybrid that was derived from the gte-large sentence transformer to extract information from unstructured clinical narratives. Our model tuning strategy for med-gte-hybrid combines contrastive learning and a denoising autoencoder. To evaluate the performance of med-gte-hybrid, we investigate several clinical prediction tasks in large patient cohorts extracted from the MIMIC-IV dataset, including Chronic Kidney Disease (CKD) patient prognosis, estimated glomerular filtration rate (eGFR) prediction, and patient mortality prediction. Furthermore, we demonstrate that the med-gte-hybrid model improves patient stratification, clustering, and text retrieval, thus outperforms current state-of-the-art models on the Massive Text Embedding Benchmark (MTEB). While some of our evaluations focus on CKD, our hybrid tuning of sentence transformers could be transferred to other medical domains and has the potential to improve clinical decision-making and personalised treatment pathways in various healthcare applications.

医学文本嵌入模型临床预测深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。