用少量数据高效发现基因互作,1%样本达全量效果
Weighted Diversified Sampling for Efficient Data-Driven Single-Cell Gene-Gene Interaction Discovery
- 两遍扫描计算样本多样性,快速生成高代表性子集
- 仅用1%单细胞数据,性能媲美全量数据训练
- 适合大规模基因互作挖掘,尤其数据受限场景
基因-基因互作在复杂人类疾病表现中起关键作用,但其发现极具挑战。本文提出一种基于数据驱动的计算方法,利用先进的Transformer模型挖掘显著基因互作。尽管Transformer模型有效,其参数密集性限制了数据效率。为此,我们引入一种新型加权多样化采样算法,仅需两次遍历数据集即可计算每个样本的多样性得分,从而高效生成用于互作发现的子集。大量实验表明,仅使用1%的单细胞数据,性能即达到使用完整数据集的效果。
原文摘要 · Abstract (English)
Gene-gene interactions play a crucial role in the manifestation of complex human diseases. Uncovering significant gene-gene interactions is a challenging task. Here, we present an innovative approach utilizing data-driven computational tools, leveraging an advanced Transformer model, to unearth noteworthy gene-gene interactions. Despite the efficacy of Transformer models, their parameter intensity presents a bottleneck in data ingestion, hindering data efficiency. To mitigate this, we introduce a novel weighted diversified sampling algorithm. This algorithm computes the diversity score of each data sample in just two passes of the dataset, facilitating efficient subset generation for interaction discovery. Our extensive experimentation demonstrates that by sampling a mere 1\% of the single-cell dataset, we achieve performance comparable to that of utilizing the entire dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。