arXiv:2410.14433q-bio.GNcs.LG2024-10

用机器学习找胆囊癌关键基因和通路,精准度高。

A Bioinformatic Approach Validated Utilizing Machine Learning Algorithms to Identify Relevant Biomarkers and Crucial Pathways in Gallbladder Cancer

  • 结合基因表达数据与机器学习筛选差异基因
  • 11个核心基因在诊断中表现最优,尤其是SLIT3等
  • 适合癌症机制研究者和生物标志物开发人员

胆囊癌(GBC)是胆道系统最常见的恶性肿瘤。明确其分子机制和生物标志物仍是研究难点。本研究整合两个来自NCBI GEO数据库的微阵列数据集(GSE100363、GSE139682),比较肿瘤与正常组织,鉴定出146个差异表达基因(39个上调,107个下调)。通过DAVID进行功能富集分析,利用STRING构建蛋白质互作网络,并应用三种中心性算法(度、最大邻近中心性、接近中心性)筛选出11个枢纽基因。同时采用皮尔逊相关性和递归特征消除法确定显著基因子集。在GSE100363数据集上训练支持向量机(SVM)和随机森林(RF)模型,并在GSE139682上验证,结果表明枢纽基因组合的分类性能最佳。最终确定NTRK2、COL14A1、SCN4B、ATP1A2、SLC17A7、SLIT3、COL7A1、CLDN4、CLEC3B、ADCYAP1R1、MFAP4为关键基因,其中SLIT3、COL7A1、CLDN4与GBC发生和发展密切相关。

原文摘要 · Abstract (English)

Gallbladder cancer (GBC) is the most frequent cause of disease among biliary tract neoplasms. Identifying the molecular mechanisms and biomarkers linked to GBC progression has been a significant challenge in scientific research. Few recent studies have explored the roles of biomarkers in GBC. Our study aimed to identify biomarkers in GBC using machine learning (ML) and bioinformatics techniques. We compared GBC tumor samples with normal samples to identify differentially expressed genes (DEGs) from two microarray datasets (GSE100363, GSE139682) obtained from the NCBI GEO database. A total of 146 DEGs were found, with 39 up-regulated and 107 down-regulated genes. Functional enrichment analysis of these DEGs was performed using Gene Ontology (GO) terms and REACTOME pathways through DAVID. The protein-protein interaction network was constructed using the STRING database. To identify hub genes, we applied three ranking algorithms: Degree, MNC, and Closeness Centrality. The intersection of hub genes from these algorithms yielded 11 hub genes. Simultaneously, two feature selection methods (Pearson correlation and recursive feature elimination) were used to identify significant gene subsets. We then developed ML models using SVM and RF on the GSE100363 dataset, with validation on GSE139682, to determine the gene subset that best distinguishes GBC samples. The hub genes outperformed the other gene subsets. Finally, NTRK2, COL14A1, SCN4B, ATP1A2, SLC17A7, SLIT3, COL7A1, CLDN4, CLEC3B, ADCYAP1R1, and MFAP4 were identified as crucial genes, with SLIT3, COL7A1, and CLDN4 being strongly linked to GBC development and prediction.

胆囊癌生物标志物机器学习基因网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。