用一次选优替代反复试错,加速基因干扰模型训练。
Efficient Data Selection for Training Genomic Perturbation Models
- 基于知识图谱设计一次性筛选干预实验的算法。
- 相比主动学习提速数月,预测准确率相当且更稳定。
- 适合需要快速构建基因调控模型的研究者使用。
基因组研究面临巨大的假设空间,而基因干扰等干预实验成本高、耗时长。为加速实验,基因干扰模型可预测干预后的转录结果。由于构建训练集困难,常采用“实验室闭环”的主动学习策略。但该方法难以利用Perturb-seq实验的天然并行性,效率低下。本文聚焦图神经网络类基因干扰模型,提出一种一次性子集选择方法:通过最大化监督信号在知识图谱中的传播范围来选取最优训练扰动。该准则基于输入知识图谱,采用次模最大化优化,保证近似最优。多数据集实验证明,该方法相比主动学习可节省数月时间,同时提升扰动选择稳定性,且保持竞争力的预测精度。
原文摘要 · Abstract (English)
Genomic studies face a vast hypothesis space, while interventions such as gene perturbations remain costly and time-consuming. To accelerate such experiments, gene perturbation models predict the transcriptional outcome of interventions. Since constructing the training set is challenging, active learning is often employed in a "lab-in-the-loop" process. While this strategy makes training more targeted, it is substantially slower, as it fails to exploit the inherent parallelizability of Perturb-seq experiments. Here, we focus on graph neural network-based gene perturbation models and propose a subset selection method that, unlike active learning, selects the training perturbations in one shot. Our method chooses the interventions that maximize the propagation of the supervision signal to the model. The selection criterion is defined over the input knowledge graph and is optimized with submodular maximization, ensuring a near-optimal guarantee. Experimental results across multiple datasets show that, in addition to providing months of acceleration compared to active learning, the method improves the stability of perturbation choices while maintaining competitive predictive accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。