arXiv:2511.22735cs.LG2025-11

整合基因和蛋白数据,预测肺癌细胞放疗敏感性。

Integrated Transcriptomic-proteomic Biomarker Identification for Radiation Response Prediction in Non-small Cell Lung Cancer Cell Lines

  • 用基因组与蛋白质组数据联合建模,筛选放疗响应标志物。
  • 融合模型在两组数据上预测准确率最高(R2=0.604)。
  • 发现的生物标志物兼具调控机制和临床应用潜力。

为构建整合转录组-蛋白质组的框架,以预测非小细胞肺癌(NSCLC)细胞系对辐射的反应,评估2 Gy照射后的存活分数(SF2)。分别从73株和46株NSCLC细胞系获取RNA测序(RNA-seq)和数据非依赖性采集质谱(DIA-MS)蛋白质组数据。预处理后保留1,605个共有的基因用于分析。采用基于频率排名的最小绝对收缩与选择算子(Lasso)回归,在五折交叉验证重复十次的条件下进行特征选择。构建仅转录组、仅蛋白质组及联合转录-蛋白质组特征集的支持向量回归(SVR)模型。通过决定系数(R²)和均方根误差(RMSE)评估模型性能。相关性分析评估了RNA与蛋白表达的一致性及关键标志物与SF2的关系。结果显示RNA与蛋白表达存在显著正相关(中位数Pearson's r = 0.363)。独立分析流程从转录组、蛋白质组及联合数据集中分别识别出20个优先基因签名。单组学模型跨组学泛化能力有限,而联合模型在两组数据中均表现出均衡的预测精度(转录组:R²=0.461,RMSE=0.120;蛋白质组:R²=0.604,RMSE=0.111)。本研究首次建立可用于NSCLC中SF2预测的蛋白-转录组框架,凸显整合两者数据的互补价值。所识别的共现生物标志物同时反映转录调控与功能性蛋白活动,提供机制洞察与转化潜力。

原文摘要 · Abstract (English)

To develop an integrated transcriptome-proteome framework for identifying concurrent biomarkers predictive of radiation response, as measured by survival fraction at 2 Gy (SF2), in non-small cell lung cancer (NSCLC) cell lines. RNA sequencing (RNA-seq) and data-independent acquisition mass spectrometry (DIA-MS) proteomic data were collected from 73 and 46 NSCLC cell lines, respectively. Following preprocessing, 1,605 shared genes were retained for analysis. Feature selection was performed using least absolute shrinkage and selection operator (Lasso) regression with a frequency-based ranking criterion under five-fold cross-validation repeated ten times. Support vector regression (SVR) models were constructed using transcriptome-only, proteome-only, and combined transcriptome-proteome feature sets. Model performance was assessed by the coefficient of determination (R2) and root mean square error (RMSE). Correlation analyses evaluated concordance between RNA and protein expression and the relationships of selected biomarkers with SF2. RNA-protein expression exhibited significant positive correlations (median Pearson's r = 0.363). Independent pipelines identified 20 prioritized gene signatures from transcriptomic, proteomic, and combined datasets. Models trained on single-omic features achieved limited cross-omic generalizability, while the combined model demonstrated balanced predictive accuracy in both datasets (R2=0.461, RMSE=0.120 for transcriptome; R2=0.604, RMSE=0.111 for proteome). This study presents the first proteotranscriptomic framework for SF2 prediction in NSCLC, highlighting the complementary value of integrating transcriptomic and proteomic data. The identified concurrent biomarkers capture both transcriptional regulation and functional protein activity, offering mechanistic insights and translational potential.

癌症预测多组学放疗敏感性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。