arXiv:2504.03847q-bio.QMcs.LG2025-04被引 5

用多模态机器学习预测肿瘤蛋白-金属结合,提升抗癌药物设计可解释性。

Interpretable Multimodal Learning for Tumor Protein-Metal Binding: Progress, Challenges, and Perspectives

  • 融合结构、序列等多模态数据构建肿瘤特异性蛋白-金属结合模型。
  • 提出数据预处理与多源信息整合策略,增强模型预测能力。
  • 聚焦可解释性,助力癌症研究中金属药物设计的可信决策。

在癌症治疗中,蛋白-金属结合机制直接影响药物的药代动力学和靶向效果,是合理设计抗肿瘤金属药物的基础。传统实验方法成本高、通量低,难以捕捉动态生物过程,而机器学习(ML)提供了有前景的替代方案。然而,当前用于肿瘤蛋白-金属结合的机器学习应用仍受限,主要挑战包括高质量肿瘤特异性数据集匮乏、多模态数据利用不足,以及复杂模型的“黑箱”特性带来的解释困难。本文综述了近年来利用机器学习预测肿瘤蛋白-金属结合的进展与挑战,涵盖数据、建模与可解释性三方面。我们介绍了多模态蛋白-金属结合数据集,并阐述了数据获取、清洗与预处理策略以支持模型训练。进一步探讨了不同数据模态间的互补价值及其整合方法。同时回顾了提升模型可解释性的技术,以增强癌症研究中决策的可信度。最后,展望了研究机遇,提出应对肿瘤蛋白数据稀缺与预测模型不足的策略,并强调两个有前景的方向:整合蛋白-蛋白相互作用数据以揭示金属结合的结构机制,以及预测金属结合后肿瘤蛋白的结构变化。

原文摘要 · Abstract (English)

In cancer therapeutics, protein-metal binding mechanisms critically govern the pharmacokinetics and targeting efficacy of drugs, thereby fundamentally shaping the rational design of anticancer metallodrugs. While conventional laboratory methods used to study such mechanisms are often costly, low throughput, and limited in capturing dynamic biological processes, machine learning (ML) has emerged as a promising alternative. Despite increasing efforts to develop protein-metal binding datasets and ML algorithms, the application of ML in tumor protein-metal binding remains limited. Key challenges include a shortage of high-quality, tumor-specific datasets, insufficient consideration of multiple data modalities, and the complexity of interpreting results due to the ''black box'' nature of complex ML models. This paper summarizes recent progress and ongoing challenges in using ML to predict tumor protein-metal binding, focusing on data, modeling, and interpretability. We present multimodal protein-metal binding datasets and outline strategies for acquiring, curating, and preprocessing them for training ML models. Moreover, we explore the complementary value provided by different data modalities and examine methods for their integration. We also review approaches for improving model interpretability to support more trustworthy decisions in cancer research. Finally, we offer our perspective on research opportunities and propose strategies to address the scarcity of tumor protein data and the limited number of predictive models for tumor protein-metal binding. We also highlight two promising directions for effective metal-based drug design: integrating protein-protein interaction data to provide structural insights into metal-binding events and predicting structural changes in tumor proteins after metal binding.

机器学习肿瘤研究蛋白-金属结合可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。