融合两种方法,选出更少更准的基因特征用于分类
BoMGene: Integrating Boruta-mRMR feature selection for enhanced Gene expression classification
- 结合Boruta与mRMR算法,优化特征选择过程
- 在25个数据集上减少特征数,同时保持或提升分类准确率
- 适合需要高效精准基因分析的研究人员使用
特征选择在分析基因表达数据中至关重要,可提升分类性能并降低高维数据的计算成本。本文提出BoMGene,一种融合Boruta与最小冗余最大相关性(mRMR)的混合特征选择方法,旨在优化特征空间并提高分类准确率。在25个公开基因表达数据集上,采用支持向量机(SVM)、随机森林、XGBoost(XGB)和梯度提升机(GBM)等常用分类器进行实验。结果表明,与仅使用mRMR相比,Boruta-mRMR组合显著减少了所选特征数量,加速了训练时间,同时保持或提升了分类准确率。该方法在多类基因表达数据分析中展现出更高的准确性、稳定性和实用性。
原文摘要 · Abstract (English)
Feature selection is a crucial step in analyzing gene expression data, enhancing classification performance, and reducing computational costs for high-dimensional datasets. This paper proposes BoMGene, a hybrid feature selection method that effectively integrates two popular techniques: Boruta and Minimum Redundancy Maximum Relevance (mRMR). The method aims to optimize the feature space and enhance classification accuracy. Experiments were conducted on 25 publicly available gene expression datasets, employing widely used classifiers such as Support Vector Machine (SVM), Random Forest, XGBoost (XGB), and Gradient Boosting Machine (GBM). The results show that using the Boruta-mRMR combination cuts down the number of features chosen compared to just using mRMR, which helps to speed up training time while keeping or even improving classification accuracy compared to using individual feature selection methods. The proposed approach demonstrates clear advantages in accuracy, stability, and practical applicability for multi-class gene expression data analysis
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。