融合蛋白序列与表达数据,提升乳腺癌分型与预后预测精度
Integrating Protein Sequence and Expression Level to Analysis Molecular Characterization of Breast Cancer Subtypes
- 用ProtGPT2生成蛋白嵌入,结合表达水平构建生物特征表示
- 对生存和生物标志物预测的F1得分分别达0.88和0.87
- 识别出KMT2C等关键蛋白,助力激素敏感与三阴性乳腺癌研究
乳腺癌的复杂性和异质性给理解其进展及指导有效治疗带来重大挑战。本研究旨在整合蛋白序列数据与表达水平,以改进乳腺癌亚型的分子表征并预测临床结局。利用专为蛋白序列设计的ProtGPT2语言模型,生成捕捉蛋白功能与结构特性的嵌入表示,并将其与蛋白表达水平融合,形成增强的生物学表征。采用集成K-means聚类与XGBoost分类等机器学习方法进行分析,成功将患者划分为具有生物学差异的群体,并准确预测了生存率与生物标志物状态,表现优异,其中生存预测的F1得分为0.88,生物标志物状态预测为0.87。特征重要性分析揭示了KMT2C、CLASP2和MYO1B在激素信号传导、细胞骨架重塑及治疗抵抗中的关键作用,可能影响激素受体阳性与三阴性乳腺癌的行为与进展。此外,蛋白-蛋白相互作用网络与相关性分析揭示了蛋白间的功能依赖关系,可能调控乳腺癌亚型的特性与演变。这些发现表明,整合蛋白序列与表达数据可深入揭示肿瘤生物学机制,显著推动个性化治疗策略的发展。
原文摘要 · Abstract (English)
Breast cancer's complexity and variability pose significant challenges in understanding its progression and guiding effective treatment. This study aims to integrate protein sequence data with expression levels to improve the molecular characterization of breast cancer subtypes and predict clinical outcomes. Using ProtGPT2, a language model specifically designed for protein sequences, we generated embeddings that capture the functional and structural properties of proteins. These embeddings were integrated with protein expression levels to form enriched biological representations, which were analyzed using machine learning methods, such as ensemble K-means for clustering and XGBoost for classification. Our approach enabled the successful clustering of patients into biologically distinct groups and accurately predicted clinical outcomes such as survival and biomarker status, achieving high performance metrics, notably an F1 score of 0.88 for survival and 0.87 for biomarker status prediction. Feature importance analysis identified KMT2C, CLASP2, and MYO1B as key proteins involved in hormone signaling, cytoskeletal remodeling, and therapy resistance in hormone receptor-positive and triple-negative breast cancer, with potential influence on breast cancer subtype behavior and progression. Furthermore, protein-protein interaction networks and correlation analyses revealed functional interdependencies among proteins that may influence the behavior and progression of breast cancer subtypes. These findings suggest that integrating protein sequence and expression data provides valuable insights into tumor biology and has significant potential to enhance personalized treatment strategies in breast cancer care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。