arXiv:2506.00053q-bio.QMcs.AI2025-06

通过无放回特征选择与随机投影提升癌症基因表达分类准确率

Improving statistical learning methods via features selection without replacement sampling and random projection

  • 结合无放回特征选择与LDA投影,降低高维数据过拟合风险
  • 从54,675个基因中筛选出20,890个关键基因,测试准确率达96%
  • 适合生物信息学研究者用于癌症基因标志物挖掘

癌症本质上是遗传疾病,由基因和表观遗传改变导致基因表达紊乱,引发细胞失控增殖与转移。高维微阵列数据因

原文摘要 · Abstract (English)

Cancer is fundamentally a genetic disease characterized by genetic and epigenetic alterations that disrupt normal gene expression, leading to uncontrolled cell growth and metastasis. High-dimensional microarray datasets pose challenges for classification models due to the "small n, large p" problem, resulting in overfitting. This study makes three different key contributions: 1) we propose a machine learning-based approach integrating the Feature Selection Without Re-placement (FSWOR) technique and a projection method to improve classification accuracy. 2) We apply the Kendall statistical test to identify the most significant genes from the brain cancer mi-croarray dataset (GSE50161), reducing the feature space from 54,675 to 20,890 genes.3) we apply machine learning models using k-fold cross validation techniques in which our model incorpo-rates ensemble classifiers with LDA projection and Naïve Bayes, achieving a test score of 96%, outperforming existing methods by 9.09%. The results demonstrate the effectiveness of our ap-proach in high-dimensional gene expression analysis, improving classification accuracy while mitigating overfitting. This study contributes to cancer biomarker discovery, offering a robust computational method for analyzing microarray data.

基因表达分类模型特征选择癌症研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。