arXiv:2501.06805cs.LGq-bio.GN2025-01被引 6

基于多视图特征选择与集成模型,精准区分33种癌症类型。

A Pan-cancer Classification Model using Multi-view Feature Selection Method and Ensemble Classifier

  • 分类型垂直分割转录组数据,多轮Boruta筛选后融合特征
  • 97.11%准确率,0.9996 AUC,12类难区分癌种识别超90%
  • 适合高维基因数据分类,尤其适用于相似组织来源癌症

精准识别癌症样本对精确诊断和有效治疗至关重要。传统方法在高维、特征数远超样本数的数据上表现不佳。本研究提出一种专用于转录组数据的新型特征选择框架,并构建两种集成分类器。首先按特征类型纵向划分转录组数据,对每部分应用Boruta特征选择,合并结果后再进行一次Boruta筛选;重复不同参数设置后生成最终特征集。随后基于LR、SVM和XGBoost构建两种集成模型,采用最大投票与概率平均策略。通过10折交叉验证评估性能。结果表明,该方法在33种癌症分类中达到97.11%准确率和0.9996 AUC,优于现有最优方法。对于12种因组织来源相似而难以区分的癌症类型,识别准确率超过90%,显著超越已有文献报道方法。基因集富集分析显示,所选特征显著富集于与癌症高度相关的通路。该框架可有效筛选与癌症发展相关特征,提升癌症类型识别精度。

原文摘要 · Abstract (English)

Accurately identifying cancer samples is crucial for precise diagnosis and effective patient treatment. Traditional methods falter with high-dimensional and high feature-to-sample count ratios, which are critical for classifying cancer samples. This study aims to develop a novel feature selection framework specifically for transcriptome data and propose two ensemble classifiers. For feature selection, we partition the transcriptome dataset vertically based on feature types. Then apply the Boruta feature selection process on each of the partitions, combine the results, and apply Boruta again on the combined result. We repeat the process with different parameters of Boruta and prepare the final feature set. Finally, we constructed two ensemble ML models based on LR, SVM and XGBoost classifiers with max voting and averaging probability approach. We used 10-fold cross-validation to ensure robust and reliable classification performance. With 97.11\% accuracy and 0.9996 AUC value, our approach performs better compared to existing state-of-the-art methods to classify 33 types of cancers. A set of 12 types of cancer is traditionally challenging to differentiate between each other due to their similarity in tissue of origin. Our method accurately identifies over 90\% of samples from these 12 types of cancers, which outperforms all known methods presented in existing literature. The gene set enrichment analysis reveals that our framework's selected features have enriched the pathways highly related to cancers. This study develops a feature selection framework to select features highly related to cancer development and leads to identifying different types of cancer samples with higher accuracy.

癌症分类特征选择集成学习转录组

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。