arXiv:2501.18741cs.LGcs.AI2025-01被引 13

小样本医疗数据用合成数据增强,能显著提升模型预测性能。

Synthetic Data Generation for Augmenting Small Samples

  • 用合成数据扩充小样本,提升数据多样性与泛化能力。
  • 增广后AUC平均提升15.55%,最高达43.23%(从0.51升至0.73)。
  • 提供决策工具,帮助判断何时适合使用数据增广。

健康研究中常面临小样本问题,导致机器学习模型泛化性能不佳。数据增强可通过增加样本量并提升数据多样性,改善模型在未见数据上的表现。研究表明,当数据集观察数少、基线AUC低、类别变量基数高且结果变量更均衡时,增强效果更显著。不同生成模型无一致优势。本文开发了决策支持模型,可判断增广是否有效。在七个真实小样本数据集上,增广使AUC提升4.31%(从0.71到0.75)至43.23%(从0.51到0.73),平均相对提升15.55%(p=0.0078)。增广组AUC高于仅重采样组(p=0.016),且数据多样性更高(p=0.046)。

原文摘要 · Abstract (English)

Small datasets are common in health research. However, the generalization performance of machine learning models is suboptimal when the training datasets are small. To address this, data augmentation is one solution. Augmentation increases sample size and is seen as a form of regularization that increases the diversity of small datasets, leading them to perform better on unseen data. We found that augmentation improves prognostic performance for datasets that: have fewer observations, with smaller baseline AUC, have higher cardinality categorical variables, and have more balanced outcome variables. No specific generative model consistently outperformed the others. We developed a decision support model that can be used to inform analysts if augmentation would be useful. For seven small application datasets, augmenting the existing data results in an increase in AUC between 4.31% (AUC from 0.71 to 0.75) and 43.23% (AUC from 0.51 to 0.73), with an average 15.55% relative improvement, demonstrating the nontrivial impact of augmentation on small datasets (p=0.0078). Augmentation AUC was higher than resampling only AUC (p=0.016). The diversity of augmented datasets was higher than the diversity of resampled datasets (p=0.046).

数据增强小样本医疗AI合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。