arXiv:2508.02743q-bio.GNcs.AI2025-08被引 1

用生成模型扩充癌症基因数据,提升分类准确率。

A Novel cVAE-Augmented Deep Learning Framework for Pan-Cancer RNA-Seq Classification

  • 用条件变分自编码器生成癌症类型相关的合成基因数据
  • 在801个样本上实现约98%的分类准确率,显著优于原始数据
  • 特别适合小样本癌症类别的分类任务

利用转录组(RNA-Seq)数据进行泛癌分类有助于肿瘤分型和治疗选择,但因维度极高且样本量有限而具挑战性。本研究提出一种新型深度学习框架,采用类别条件变分自编码器(cVAE)对泛癌基因表达分类的训练数据进行增强。基于来自癌症基因组图谱(TCGA)的801个肿瘤RNA-Seq样本(涵盖5种癌症类型),首先通过特征选择将20,531个基因表达特征缩减至500个变异度最高的基因。随后在该数据上训练cVAE,学习以癌症类型为条件的基因表达潜在表示,从而生成每类肿瘤的合成基因表达样本。将生成样本加入训练集(使数据量翻倍),以缓解过拟合与类别不平衡问题。再使用两层多层感知机(MLP)分类器在增广数据集上训练,对肿瘤类型进行预测。增广框架在独立测试集上达到约98%的分类准确率,显著优于仅使用原始数据训练的分类器。实验包含VAE训练曲线、分类器性能指标(ROC曲线与混淆矩阵)及架构图,结果表明,基于cVAE的合成数据增强可显著提升泛癌预测性能,尤其对样本较少的癌症类别效果明显。

原文摘要 · Abstract (English)

Pan-cancer classification using transcriptomic (RNA-Seq) data can inform tumor subtyping and therapy selection, but is challenging due to extremely high dimensionality and limited sample sizes. In this study, we propose a novel deep learning framework that uses a class-conditional variational autoencoder (cVAE) to augment training data for pan-cancer gene expression classification. Using 801 tumor RNA-Seq samples spanning 5 cancer types from The Cancer Genome Atlas (TCGA), we first perform feature selection to reduce 20,531 gene expression features to the 500 most variably expressed genes. A cVAE is then trained on this data to learn a latent representation of gene expression conditioned on cancer type, enabling the generation of synthetic gene expression samples for each tumor class. We augment the training set with these cVAE-generated samples (doubling the dataset size) to mitigate overfitting and class imbalance. A two-layer multilayer perceptron (MLP) classifier is subsequently trained on the augmented dataset to predict tumor type. The augmented framework achieves high classification accuracy (~98%) on a held-out test set, substantially outperforming a classifier trained on the original data alone. We present detailed experimental results, including VAE training curves, classifier performance metrics (ROC curves and confusion matrix), and architecture diagrams to illustrate the approach. The results demonstrate that cVAE-based synthetic augmentation can significantly improve pan-cancer prediction performance, especially for underrepresented cancer classes.

泛癌分类生成模型数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。