用遗传算法生成指定难度的合成数据集,提升机器学习评估的多样性。
Transforming Datasets to Requested Complexity with Projection-based Many-Objective Genetic Algorithm
- 基于多目标遗传算法,通过线性投影调节数据复杂度。
- 可精准生成符合目标复杂度的分类与回归数据集。
- 适合需要可控难度数据的模型评测与算法开发人员。
研究社区持续寻求更先进的合成数据生成方法,以可靠评估机器学习方法的优势与局限。本文提出一种遗传算法,通过优化分类与回归任务中的问题复杂度指标,使合成数据集达到特定目标复杂度。针对分类任务,采用10个复杂度度量;针对回归任务,选取4个表现出良好优化能力的度量。实验表明,该算法可通过线性特征投影,将合成数据集转化为具有不同难度的目标复杂度。对前沿分类器和回归器的评估显示,生成数据的复杂度与识别性能之间存在显著相关性。
原文摘要 · Abstract (English)
The research community continues to seek increasingly more advanced synthetic data generators to reliably evaluate the strengths and limitations of machine learning methods. This work aims to increase the availability of datasets encompassing a diverse range of problem complexities by proposing a genetic algorithm that optimizes a set of problem complexity measures for classification and regression tasks towards specific targets. For classification, a set of 10 complexity measures was used, while for regression tasks, 4 measures demonstrating promising optimization capabilities were selected. Experiments confirmed that the proposed genetic algorithm can generate datasets with varying levels of difficulty by transforming synthetically created datasets to achieve target complexity values through linear feature projections. Evaluations involving state-of-the-art classifiers and regressors revealed a correlation between the complexity of the generated data and the recognition quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。