融合量子特征与多模态数据,提升酶功能分类准确率
Multimodal Quantum Vision Transformer for Enzyme Commission Classification from Biochemical Representations
- 用四种生化模态融合建模:序列、量子电子描述符、分子图和2D图像
- 在酶分类任务上达85.1%准确率,显著优于仅用序列的模型
- 适合对量子机器学习与生物信息交叉研究感兴趣的读者
准确预测酶功能仍是计算生物学中的重大挑战,尤其针对结构注释或序列同源性有限的酶。本文提出一种新型多模态量子机器学习(QML)框架,通过整合四种互补的生化模态——蛋白序列嵌入、量子衍生的电子描述符、分子图结构和2D分子图像表示——来增强酶委员会(EC)分类。模型采用量子视觉变换器(QVT)主干网络,配备模态专用编码器和统一的跨注意力融合模块。通过融合图结构特征与空间模式,该方法捕捉了酶功能背后的立体电子相互作用。实验表明,所提多模态QVT模型在顶1准确率达85.1%,显著优于仅使用序列的基线模型,并在与其他QML模型对比中表现更优。
原文摘要 · Abstract (English)
Accurately predicting enzyme functionality remains one of the major challenges in computational biology, particularly for enzymes with limited structural annotations or sequence homology. We present a novel multimodal Quantum Machine Learning (QML) framework that enhances Enzyme Commission (EC) classification by integrating four complementary biochemical modalities: protein sequence embeddings, quantum-derived electronic descriptors, molecular graph structures, and 2D molecular image representations. Quantum Vision Transformer (QVT) backbone equipped with modality-specific encoders and a unified cross-attention fusion module. By integrating graph features and spatial patterns, our method captures key stereoelectronic interactions behind enzyme function. Experimental results demonstrate that our multimodal QVT model achieves a top-1 accuracy of 85.1%, outperforming sequence-only baselines by a substantial margin and achieving better performance results compared to other QML models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。