用机器学习自动分类罕见且难辨的冷球蛋白病,减少对专家判断的依赖。
Machine Learning Classification of Cryopathy Syndromes: A Comprehensive Comparative Study
- 构建12种模型对比,用临床特征和交互特征提升分类能力
- 集成模型在14类诊断中达到最高宏F1分数,稳定表现优于神经网络
- 适合需要辅助诊断冷球蛋白病的临床医生和研究者参考
冷球蛋白病综合征难以分类,因实验室指标常重叠且部分诊断罕见,导致常规检测解读困难,依赖专家判断。本研究旨在开发并比较基于实验室数据的机器学习方法,实现冷球蛋白病综合征的自动化分类,为临床决策提供支持。分析了2686名患者的实验室记录,涵盖14个诊断类别,包括人口学变量、冷球蛋白测定、沉淀试验以及血凝素和溶血素滴度。数据预处理包括清洗、编码、缺失值填补、归一化及构建临床相关交互特征。评估了12种建模策略:随机森林、梯度提升树、多层感知机、软投票集成、合成少数类过采样技术(SMOTE)进行类别平衡、层次分类、周期感知模型、目标二分类器及概率校准。性能通过分层训练-测试划分和分层5折交叉验证评估,主要指标为宏平均F1分数、准确率、前3准确率和预期校准误差。整体任务因显著类别不平衡和诊断间临床重叠而困难。最佳多分类性能由随机森林与梯度提升树的软投票集成实现。交叉验证证实平衡随机森林模型表现稳定。树基方法始终优于神经网络模型。特征工程提升了区分能力,最具信息量的预测因子为基于冷球蛋白的交互特征。
原文摘要 · Abstract (English)
Cryopathy syndromes are difficult to classify because laboratory patterns often overlap across diagnostic categories, while some diagnoses are rare. This makes routine interpretation of cryoglobulin-related tests challenging and increases dependence on expert judgment. The aim of this study was to develop and compare machine learning approaches for automated classification of cryopathy syndromes from laboratory data and to identify a practical strategy for clinical decision support. Methods: We analysed laboratory records from 2,686 patients assigned to 14 diagnostic categories. The dataset included demographic variables, cryoglobulin measurements, precipitation tests, and hemagglutinin and hemolysin titers. Data preprocessing included cleaning, encoding, imputation, normalization, and construction of clinically informed interaction features. We evaluated 12 modelling strategies, including Random Forest, Gradient Boosted Trees, Multi-Layer Perceptron, soft-voting ensembles, class balancing with Synthetic Minority Over-sampling Technique, hierarchical classification, period-aware models, targeted binary classifiers, and probability calibration. Performance was assessed using stratified train-test evaluation and stratified 5-fold cross-validation. The main metrics were macro-averaged F1 score, accuracy, Top-3 accuracy, and expected calibration error. The overall task proved difficult because of marked class imbalance and clinical overlap between diagnoses. The best multiclass performance was achieved by a soft-voting ensemble of Random Forest and Gradient Boosted Trees. Cross-validation confirmed stable performance for the balanced Random Forest model. Tree-based methods consistently outperformed the neural network model. Feature engineering improved discrimination, and the most informative predictors were derived cryoglobulin-based interaction features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。