arXiv:2411.11436cs.LGcs.AI2024-11

用隐式正则化实现多标签特征选择,减少偏差并避免过拟合。

Implicit Regularization for Multi-label Feature Selection

  • 通过哈达玛积参数化实现隐式正则化,无需显式惩罚项。
  • 在多个基准数据集上表现更优,偏差更小且可能产生良性过拟合。
  • 适合多标签学习中追求低偏差与鲁棒性的研究者。

本文针对多标签学习中的特征选择问题,提出一种基于隐式正则化与标签嵌入的新估计器。与使用 $l_{2,1}$-范数、MCP 或 SCAD 等显式正则化项的稀疏方法不同,本文采用哈达玛积参数化提供一种简洁替代方案。为引导特征选择过程,引入多标签信息的潜在语义表示作为标签嵌入。在若干已知基准数据集上的实验表明,所提估计器具有更小的额外偏差,可能引发良性过拟合。

原文摘要 · Abstract (English)

In this paper, we address the problem of feature selection in the context of multi-label learning, by using a new estimator based on implicit regularization and label embedding. Unlike the sparse feature selection methods that use a penalized estimator with explicit regularization terms such as $l_{2,1}$-norm, MCP or SCAD, we propose a simple alternative method via Hadamard product parameterization. In order to guide the feature selection process, a latent semantic of multi-label information method is adopted, as a label embedding. Experimental results on some known benchmark datasets suggest that the proposed estimator suffers much less from extra bias, and may lead to benign overfitting.

特征选择多标签学习隐式正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。