通过最小化邻域Jensen散度提升符号回归的泛化能力
Enhancing Generalization in Evolutionary Feature Construction for Symbolic Regression through Vicinal Jensen Gap Minimization
- 用邻域Jensen散度替代传统复杂度度量,指导特征构造进化
- 在58个数据集上优于15种算法,显著降低过拟合风险
- 自适应调节正则强度并检测虚假样本,适合高噪声场景
基于遗传编程的特征构造近年来在自动化机器学习中表现优异,但过拟合问题仍限制其广泛应用。本文证明,通过噪声扰动或mixup数据增强估计的邻域风险,可被经验风险与正则项(有限差分或邻域Jensen散度)之和所界定。基于此分解,提出一种联合优化经验风险与邻域Jensen散度的进化特征构造框架,以控制过拟合。针对不同数据集的噪声水平差异,设计噪声估计策略动态调整正则强度;为缓解流形侵入(即数据增强生成脱离真实数据流形的不合理样本),提出流形侵入检测机制。在58个数据集上的实验表明,相比其他复杂度度量,邻域Jensen散度最小化显著提升泛化性能;与15种主流机器学习算法对比,该方法在符号回归任务中取得更优结果。
原文摘要 · Abstract (English)
Genetic programming-based feature construction has achieved significant success in recent years as an automated machine learning technique to enhance learning performance. However, overfitting remains a challenge that limits its broader applicability. To improve generalization, we prove that vicinal risk, estimated through noise perturbation or mixup-based data augmentation, is bounded by the sum of empirical risk and a regularization term-either finite difference or the vicinal Jensen gap. Leveraging this decomposition, we propose an evolutionary feature construction framework that jointly optimizes empirical risk and the vicinal Jensen gap to control overfitting. Since datasets may vary in noise levels, we develop a noise estimation strategy to dynamically adjust regularization strength. Furthermore, to mitigate manifold intrusion-where data augmentation may generate unrealistic samples that fall outside the data manifold-we propose a manifold intrusion detection mechanism. Experimental results on 58 datasets demonstrate the effectiveness of Jensen gap minimization compared to other complexity measures. Comparisons with 15 machine learning algorithms further indicate that genetic programming with the proposed overfitting control strategy achieves superior performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。