arXiv:2504.00306cs.LGq-bio.GN2025-04被引 6

提出基因组级留一染色体验证法,避免模型性能虚高。

LOCO-EPI: Leave-one-chromosome-out (LOCO) as a benchmarking paradigm for deep learning based prediction of enhancer-promoter interactions

  • 用留一染色体法替代随机分割,防止基因组信息泄露。
  • 传统模型在留一染色体下性能大幅下降,暴露过拟合问题。
  • 新混合神经网络在严格验证下仍表现优异,适合真实场景应用。

在哺乳动物和脊椎动物基因组中,基因启动子与其远端增强子可能相距数百万碱基对,且一个启动子未必与最近的增强子相互作用。由于碱基对距离无法有效预测此类互作,研究者开发了多种机器学习方法用于预测增强子-启动子互作(EPI)。现有方法通常将EP对随机划分为训练与测试集,但这种划分方式会导致同一基因组区域的样本同时出现在训练和测试集中,造成信息泄露,高估模型性能。本文提出采用更严格的留一染色体(LOCO)交叉验证作为评估范式。实验表明,此前在随机分割下表现良好的深度学习模型在LOCO设置下性能显著下降,证实了性能被过度估计。此外,我们提出一种融合k-mer序列特征的新型混合深度神经网络,在LOCO设置下表现显著优于基线模型,说明其能捕捉更具泛化性的互作规律。本文还公开了基于LOCO划分的EPI数据集,数据可访问:https://github.com/malikmtahir/EPI。

原文摘要 · Abstract (English)

In mammalian and vertebrate genomes, the promoter regions of the gene and their distal enhancers may be located millions of base-pairs from each other, while a promoter may not interact with the closest enhancer. Since base-pair proximity is not a good indicator of these interactions, there is considerable work toward developing methods for predicting Enhancer-Promoter Interactions (EPI). Several machine learning methods have reported increasingly higher accuracies for predicting EPI. Typically, these approaches randomly split the dataset of Enhancer-Promoter (EP) pairs into training and testing subsets followed by model training. However, the aforementioned random splitting causes information leakage by assigning EP pairs from the same genomic region to both testing and training sets, leading to performance overestimation. In this paper we propose to use a more thorough training and testing paradigm i.e., Leave-one-chromosome-out (LOCO) cross-validation for EPI-prediction. We demonstrate that a deep learning algorithm, which gives higher accuracies when trained and tested on random-splitting setting, drops drastically in performance under LOCO setting, confirming overestimation of performance. We further propose a novel hybrid deep neural network for EPI-prediction that fuses k-mer features of the nucleotide sequence. We show that the hybrid architecture performs significantly better in the LOCO setting, demonstrating it can learn more generalizable aspects of EP interactions. With this paper we are also releasing the LOCO splitting-based EPI dataset. Research data is available in this public repository: https://github.com/malikmtahir/EPI

基因互作深度学习数据验证基因组学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。