融合结构化与文本数据,用机器学习模型更早预测房颤复发风险。
Early Diagnosis of Atrial Fibrillation Recurrence: A Large Tabular Model Approach with Structured and Unstructured Clinical Data
- 将电子病历文本与结构化数据结合,提升数据质量。
- 新模型在1~2年内预测房颤复发准确率优于传统评分和普通机器学习。
- 发现性别与年龄存在预测偏差,提示需关注模型公平性。
房颤是常见心律失常,与高发病率和死亡率相关。在治疗快速演进的背景下,早期预测房颤复发对优化治疗至关重要,但传统评分(如CHADS2-VASc、HATCH、APPLE)预测能力有限。以往研究多依赖编码的电子健康记录,易含错误和缺失。本研究通过自然语言处理技术整合结构化临床数据与自由文本出院报告,构建高质量表格数据集,纳入1,508例确诊房颤患者,评估传统评分、机器学习模型及提出的LTM方法。结果表明,LTM模型在手动标注测试集上表现最优,显著超越传统评分和主流机器学习模型。此外,性别与年龄分析揭示了预测偏差,凸显传统评分局限性,并验证了基于机器学习方法的潜力。
原文摘要 · Abstract (English)
BACKGROUND: Atrial fibrillation (AF), the most common arrhythmia, is linked to high morbidity and mortality. In a fast-evolving AF rhythm control treatment era, predicting AF recurrence after its onset may be crucial to achieve the optimal therapeutic approach, yet traditional scores like CHADS2-VASc, HATCH, and APPLE show limited predictive accuracy. Moreover, early diagnosis studies often rely on codified electronic health record (EHR) data, which may contain errors and missing information. OBJECTIVE: This study aims to predict AF recurrence between one month and two years after onset by evaluating traditional clinical scores, ML models, and our LTM approach. Moreover, another objective is to develop a methodology for integrating structured and unstructured data to enhance tabular dataset quality. METHODS: A tabular dataset was generated by combining structured clinical data with free-text discharge reports processed through natural language processing techniques, reducing errors and annotation effort. A total of 1,508 patients with documented AF onset were identified, and models were evaluated on a manually annotated test set. The proposed approach includes a LTM compared against traditional clinical scores and ML models. RESULTS: The proposed LTM approach achieved the highest predictive performance, surpassing both traditional clinical scores and ML models. Additionally, the gender and age bias analyses revealed demographic disparities. CONCLUSION: The integration of structured data and free-text sources resulted in a high-quality dataset. The findings emphasize the limitations of traditional clinical scores in predicting AF recurrence and highlight the potential of ML-based approaches, particularly our LTM model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。