arXiv:2505.14587stat.MLcs.LG2025-05被引 4

分析高维下自助法集成分类器的理论性能,提出优化策略。

High-Dimensional Analysis of Bootstrap Ensemble Classifiers

  • 用随机矩阵理论分析自助法在高维数据中的表现
  • 发现样本量和特征维度增大时性能变化规律
  • 给出子集数量与正则化参数的选择建议,适合机器学习研究者

自助法长期是集成学习的核心方法。本文针对大样本量和高特征维度场景,对基于最小二乘支持向量机(LSSVM)的自助集成分类器进行理论分析。利用随机矩阵理论,研究由多个弱分类器在不同数据子集上训练得到的决策函数集成后的性能。揭示了自助法在高维设置下的作用机制,深化了对其影响的理解。基于理论发现,提出了最大化LSSVM性能的子集数量与正则化参数选择策略。在合成数据与真实世界数据集上的实验验证了理论结果的有效性。

原文摘要 · Abstract (English)

Bootstrap methods have long been the cornerstone of ensemble learning in machine learning. This paper presents a theoretical analysis of bootstrap techniques applied to the Least Square Support Vector Machine (LSSVM) ensemble in the context of large and growing sample sizes and feature dimensionalities. Using tools from Random Matrix Theory, we investigate the performance of this classifier that aggregates decision functions from multiple weak classifiers, each trained on different subsets of the data. We provide insights into the use of bootstrap methods in high-dimensional settings, enhancing our understanding of their impact. Based on these findings, we propose strategies to select the number of subsets and the regularization parameter that maximize the performance of the LSSVM. Empirical experiments on synthetic and real-world datasets validate our theoretical results.

集成学习高维数据理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。