arXiv:2506.17326cs.LGstat.AP2025-06被引 7

用概率模型生成糖尿病数据,提升少数类预测效果

CopulaSMOTE: A Copula-Based Oversampling Approach for Imbalanced Classification in Diabetes Prediction

  • 基于藤蔓耦合模型捕捉糖尿病数据的复杂依赖关系
  • 在三大糖尿病数据集上,对少数类召回率提升显著
  • 适合处理高维、不平衡的医疗预测任务

类别不平衡仍是糖尿病等临床预测模型开发中的实际难题,确诊病例数远少于健康对照。传统SMOTE及其变体通过特征空间局部插值生成合成样本,但未显式建模少数类的联合依赖结构。本文提出基于耦合的过采样方法CopulaSMOTE,通过截断藤蔓耦合(truncated vine copulas)以一系列二元构建块表示多变量依赖关系,在生成合成样本时显式估计少数类结构,并与标准机器学习方法集成。我们在三个公开糖尿病数据集上评估该方法:皮马印第安人糖尿病数据集、伊拉克糖尿病数据集和美国疾控中心BRFSS 2015糖尿病健康指标数据集,涵盖不同样本量、维度和不平衡程度。每组数据中对比五种重采样策略与五种分类器,采用5×2交叉验证协议并进行Dietterich配对t检验。结果表明,CopulaSMOTE在较大表格型糖尿病数据集(特别是CDC BRFSS数据集)中能有效提升少数类识别能力,但其优势依赖于具体分类器和评估指标。

原文摘要 · Abstract (English)

Class imbalance remains a practical obstacle in the development of clinical prediction models for conditions such as diabetes mellitus, where the number of confirmed cases is often much smaller than the number of controls. The Synthetic Minority Over-sampling Technique (SMOTE) and its variants are widely used to address this imbalance, but they generate synthetic observations through local interpolation in feature space and do not explicitly model the joint dependence structure of the minority class. To address this challenge, our study introduces a copula-based data augmentation approach that estimates the minority-class dependence structure when generating synthetic samples and integrates with standard machine learning techniques. Specifically, we employ truncated vine copulas to represent multivariate dependence through a sequence of bivariate building blocks. We evaluate the proposed approach on three public diabetes datasets, namely the Pima Indians Diabetes dataset, the Iraqi Diabetes dataset, and the CDC BRFSS 2015 Diabetes Health Indicators dataset, which together cover a range of sample sizes, dimensionalities, and imbalance regimes. For each dataset, five resampling strategies are compared across five classifiers using a 5 by 2 cross validation protocol with Dietterich's paired t test. Our findings suggest that CopulaSMOTE can improve minority-class recovery in larger tabular diabetes datasets, particularly the CDC BRFSS dataset, but its advantages depend on the classifier and evaluation metric.

医疗预测数据增强不平衡学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。