针对高维数据,提出按类别独立选特征的新方法,提升分类性能。
Feature Selection for Latent Factor Models
- 基于低秩生成模型,为每类单独建模并选特征
- 在标准数据集上优于现有方法,理论可保证正确恢复真特征
- 适合高维分类任务中需要精准特征筛选的场景
特征选择在高维数据中至关重要,有助于识别相关特征、缓解维度灾难并提升机器学习性能。传统分类特征选择方法使用所有类的数据为每个类别选取特征。本文探索了为每类分别建模的方法,基于低秩生成模型,并引入信噪比(SNR)作为特征选择准则。该新方法在特定假设下具备理论上的真特征恢复保证,且在标准分类数据集上表现优于部分现有方法。
原文摘要 · Abstract (English)
Feature selection is crucial for pinpointing relevant features in high-dimensional datasets, mitigating the 'curse of dimensionality,' and enhancing machine learning performance. Traditional feature selection methods for classification use data from all classes to select features for each class. This paper explores feature selection methods that select features for each class separately, using class models based on low-rank generative methods and introducing a signal-to-noise ratio (SNR) feature selection criterion. This novel approach has theoretical true feature recovery guarantees under certain assumptions and is shown to outperform some existing feature selection methods on standard classification datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。