arXiv:2410.04774cs.LG2024-10中稿 · 05 October 2024被引 23

用粗粒度球体改进支持向量机,提升抗噪性和大规模计算效率。

Granular Ball Twin Support Vector Machine

  • 用数据球体替代单点,增强模型稳定性
  • 无需矩阵求逆,计算更快且可扩展
  • 适合噪声多、数据量大的分类任务

孪生支持向量机(TSVM)在分类与回归中应用广泛,但面临三大挑战:(i) 矩阵求逆导致大规模数据下效率低下;(ii) 原始形式未遵循结构风险最小化(SRM)原则,易过拟合;(iii) 对噪声和异常值敏感,重采样时不稳定。为此,本文提出粒度球孪生支持向量机(GBTSVM),以粒度球代替单个数据点作为输入,提升对噪声和重采样的鲁棒性。进一步提出大规模粒度球孪生支持向量机(LS-GBTSVM),其优化公式避免了矩阵求逆,显著提升计算效率,并通过正则化项引入SRM原则,有效缓解过拟合。在UCI、KEEL和NDC等基准数据集上的实验表明,所提模型具备更强的泛化能力与可扩展性。

原文摘要 · Abstract (English)

On Efficient and Scalable Computation of the Nonparametric Maximum Likelihood Estimator in Mixture ModelsTwin support vector machine (TSVM) is an emerging machine learning model with versatile applicability in classification and regression endeavors. Nevertheless, TSVM confronts noteworthy challenges: $(i)$ the imperative demand for matrix inversions presents formidable obstacles to its efficiency and applicability on large-scale datasets; $(ii)$ the omission of the structural risk minimization (SRM) principle in its primal formulation heightens the vulnerability to overfitting risks; and $(iii)$ the TSVM exhibits a high susceptibility to noise and outliers, and also demonstrates instability when subjected to resampling. In view of the aforementioned challenges, we propose the granular ball twin support vector machine (GBTSVM). GBTSVM takes granular balls, rather than individual data points, as inputs to construct a classifier. These granular balls, characterized by their coarser granularity, exhibit robustness to resampling and reduced susceptibility to the impact of noise and outliers. We further propose a novel large-scale granular ball twin support vector machine (LS-GBTSVM). LS-GBTSVM's optimization formulation ensures two critical facets: $(i)$ it eliminates the need for matrix inversions, streamlining the LS-GBTSVM's computational efficiency, and $(ii)$ it incorporates the SRM principle through the incorporation of regularization terms, effectively addressing the issue of overfitting. The proposed LS-GBTSVM exemplifies efficiency, scalability for large datasets, and robustness against noise and outliers. We conduct a comprehensive evaluation of the GBTSVM and LS-GBTSVM models on benchmark datasets from UCI, KEEL, and NDC datasets. Our experimental findings and statistical analyses affirm the superior generalization prowess of the proposed GBTSVM and LS-GBTSVM models.

支持向量机抗噪分类大规模学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。