arXiv:2410.00759cs.LGstat.ML2024-10被引 5

基于难例特征生成高价值合成数据,提升表格数据模型表现

Targeted synthetic data generation for tabular data via hardness characterization

  • 用难例特征筛选关键数据点,仅生成高价值合成样本
  • 在多个表格数据集上,合成数据使模型预测更准确
  • 相比非目标生成,计算效率更高,适合数据稀缺场景

通过合成数据生成进行数据增强已被证明能有效提升稀疏或低质量数据下的模型性能与鲁棒性。我们引入一种简单高效的增强流程,基于难例特征仅生成高价值训练样本,利用数据估值框架统计识别有益与有害观测。实证表明,基于Shapley的数据估值方法在难例识别任务中表现媲美学习型方法,且计算成本显著更低。进一步实验显示,基于最难样本训练的合成数据生成器,在多个表格数据集上优于非目标数据增强方法。该方法提升了模型外样本预测质量,同时比非目标方法更具计算效率。

原文摘要 · Abstract (English)

Data augmentation via synthetic data generation has been shown to be effective in improving model performance and robustness in the context of scarce or low-quality data. Using the data valuation framework to statistically identify beneficial and detrimental observations, we introduce a simple augmentation pipeline that generates only high-value training points based on hardness characterization, in a computationally efficient manner. We first empirically demonstrate via benchmarks on real data that Shapley-based data valuation methods perform comparably with learning-based methods in hardness characterization tasks, while offering significant computational advantages. Then, we show that synthetic data generators trained on the hardest points outperform non-targeted data augmentation on a number of tabular datasets. Our approach improves the quality of out-of-sample predictions and it is computationally more efficient compared to non-targeted methods.

数据增强合成数据表格数据难例挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。