arXiv:2603.09685cs.CLcs.AI2026-03

用大上下文电子病历自动分类老人心血管风险,效果优于传统方法。

Automatic Cardiac Risk Management Classification using large-context Electronic Patients Health Records

  • 设计专用Transformer模型处理长篇医疗文本中的复杂依赖关系。
  • 在3482例患者数据上,模型F1分数和相关系数均领先于其他方法。
  • 适合需要自动化临床风险分层的医院与研究机构使用。

为克服老年人心血管风险管理中人工编码的局限性,本研究提出一种基于非结构化电子健康记录(EHR)的自动化分类框架。基于包含3,482名患者的纵向荷兰临床文本数据集,对比了三类建模范式:经典机器学习基线、专为长序列优化的深度学习架构,以及零样本设置下的通用生成型大语言模型(LLMs)。此外,还评估了将非结构化文本与结构化药物嵌入及人体测量数据融合的后期融合策略。结果表明,自研Transformer架构在性能上超越传统方法和生成式模型,取得最高F1分数与马修斯相关系数。该发现凸显了专用层次化注意力机制在捕捉医学文本长距离依赖关系中的关键作用,为临床风险分层提供了高效可靠的自动化替代方案。

原文摘要 · Abstract (English)

To overcome the limitations of manual administrative coding in geriatric Cardiovascular Risk Management, this study introduces an automated classification framework leveraging unstructured Electronic Health Records (EHRs). Using a dataset of 3,482 patients, we benchmarked three distinct modeling paradigms on longitudinal Dutch clinical narratives: classical machine learning baselines, specialized deep learning architectures optimized for large-context sequences, and general-purpose generative Large Language Models (LLMs) in a zero-shot setting. Additionally, we evaluated a late fusion strategy to integrate unstructured text with structured medication embeddings and anthropometric data. Our analysis reveals that the custom Transformer architecture outperforms both traditional methods and generative \acs{llm}s, achieving the highest F1-scores and Matthews Correlation Coefficients. These findings underscore the critical role of specialized hierarchical attention mechanisms in capturing long-range dependencies within medical texts, presenting a robust, automated alternative to manual workflows for clinical risk stratification.

心血管风险电子病历Transformer自动化分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。