用先进数据增强提升深度学习在表格回归中的表现
Data Augmentation for Deep Learning Regression Tasks by Machine Learning Models
- 设计保留统计关系的高级数据增强策略
- 平均性能提升超10%,30个数据集验证有效
- 适合想用深度学习做表格回归的研究者
深度学习模型在计算机视觉和自然语言处理中表现突出,但在涉及表格数据的回归任务中仍应用不足,传统机器学习模型常更优。本文提出并评估多种数据增强(DA)技术,以提升深度学习模型在表格回归任务中的表现。比较了从简单复制加噪声到更复杂的、保持数据底层统计关系的增强策略。分析表明,先进增强方法在多个数据集和回归任务中显著提升深度学习模型性能,相比无增强基线平均提升超过10%。该有效性在30个不同数据集上通过三次迭代与三种自动化深度学习框架(AutoKeras、H2O、AutoGluon)严格验证。研究证明,通过采用先进数据增强技术,深度学习模型可在回归任务中发挥全部潜力,推动其在实际应用中的更广泛应用与性能提升。
原文摘要 · Abstract (English)
Deep learning (DL) models have gained prominence in domains such as computer vision and natural language processing but remain underutilized for regression tasks involving tabular data. In these cases, traditional machine learning (ML) models often outperform DL models. In this study, we propose and evaluate various data augmentation (DA) techniques to improve the performance of DL models for tabular data regression tasks. We compare the performance gain of Neural Networks by different DA strategies ranging from a naive method of duplicating existing observations and adding noise to a more sophisticated DA strategy that preserves the underlying statistical relationship in the data. Our analysis demonstrates that the advanced DA method significantly improves DL model performance across multiple datasets and regression tasks, resulting in an average performance increase of over 10\% compared to baseline models without augmentation. The efficacy of these DA strategies was rigorously validated across 30 distinct datasets, with multiple iterations and evaluations using three different automated deep learning (AutoDL) frameworks: AutoKeras, H2O, and AutoGluon. This study demonstrates that by leveraging advanced DA techniques, DL models can realize their full potential in regression tasks, thereby contributing to broader adoption and enhanced performance in practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。