arXiv:2409.02588cs.LGq-bio.BM2024-09被引 9

用多视图网络提升DNA结合蛋白预测准确率

Multiview Random Vector Functional Link Network for Predicting DNA-Binding Proteins

  • 融合三种蛋白特征,通过多视图学习增强模型表达能力
  • 在多个数据集上表现优于基线模型,泛化性能更强
  • 适合生物信息学与机器学习交叉研究者参考

DNA结合蛋白(DBPs)的识别对理解生命活动至关重要。近年来,基于机器学习的模型被广泛用于DBP预测。本文提出一种新型多视图随机向量函数链接(MvRVFL)网络框架,结合神经网络与多视图学习优势,兼顾早期与晚期融合优点。该模型为每个视图设置独立正则化参数,并采用闭式解高效求解未知参数。主目标函数包含耦合项,旨在最小化所有视图的综合误差。从三个蛋白视图中各提取五种特征,训练过程中融合隐藏特征。实验表明,MvRVFL在DBP数据集上的表现优于基线模型,且在多个基准数据集上验证了其有效性。理论分析与实证结果一致显示其具有更优泛化性能。

原文摘要 · Abstract (English)

The identification of DNA-binding proteins (DBPs) is essential due to their significant impact on various biological activities. Understanding the mechanisms underlying protein-DNA interactions is essential for elucidating various life activities. In recent years, machine learning-based models have been prominently utilized for DBP prediction. In this paper, to predict DBPs, we propose a novel framework termed a multiview random vector functional link (MvRVFL) network, which fuses neural network architecture with multiview learning. The MvRVFL model integrates both late and early fusion advantages, enabling separate regularization parameters for each view, while utilizing a closed-form solution for efficiently determining unknown parameters. The primal objective function incorporates a coupling term aimed at minimizing a composite of errors stemming from all views. From each of the three protein views of the DBP datasets, we extract five features. These features are then fused together by incorporating a hidden feature during the model training process. The performance of the proposed MvRVFL model on the DBP dataset surpasses that of baseline models, demonstrating its superior effectiveness. We further validate the practicality of the proposed model across diverse benchmark datasets, and both theoretical analysis and empirical results consistently demonstrate its superior generalization performance over baseline models.

蛋白质预测多视图学习机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。