用课程引导多模态学习,预测纳米材料与蛋白质相互作用
Curriculum-guided multimodal representation learning enables generalizable prediction of nanomaterial-protein interactions
- 分阶段课程设计,逐步扩大生物环境覆盖范围
- 在5个指标上平均得分超0.75,对未知材料/蛋白表现稳健
- 适合药物研发和纳米材料设计人员快速评估相互作用
纳米材料-蛋白质相互作用(NPI)是实现纳米材料治疗与诊断潜力的关键。尽管人工智能有望加速机制理解并推动理性设计,但对未见纳米材料或蛋白质的鲁棒泛化仍无解。本文提出CuMMI(课程引导的多模态交互模型),一个可泛化、可解释、可迁移的模型,用于复杂生物环境中预测NPI。CuMMI基于自构建的百万级NPI数据集,采用以人血浆为中心的多阶段课程,逐步扩展至更广谱生物液环境,提升数据覆盖率与泛化能力。通过融合蛋白质序列、结构及37个特征的文本编码实验背景,模型捕获材料特异性、生化及环境信息。为确保数据充分利用,对样本赋予质量权重,缓解低置信度与稀疏记录条目影响。消融实验证明关键表格特征贡献。在独立保留的时间、纳米材料、蛋白质三种外推验证中,框架持续表现优异(五项分类指标平均超过0.75),凸显其鲁棒性与泛化能力。此外,在独立金纳米颗粒数据与保留蛋白质子集上微调,仅用少量样本即优于从零训练。本方法实现了可泛化、可迁移的NPI预测,有望加速体外研究与纳米材料应用。
原文摘要 · Abstract (English)
Nanomaterial-protein interactions (NPI) are pivotal to realizing the therapeutic and diagnostic potential of nanomaterials. Although AI promises to accelerate mechanistic understanding and enable rational nanomaterial design, robust generalization to unseen nanomaterials or proteins remains unresolved. Here, we present CuMMI (curriculum-guided multimodal interaction model), a generalizable, explainable, and transferable model designed to infer NPI across complex biological settings. CuMMI leverages a self-constructed million-scale NPI dataset and adopts a multi-stage curriculum centered on human plasma, with progressively broader biofluid exposure to enhance data coverage and generalizability. By integrating protein sequence, structure, and a text-encoded experimental context of 37 features, CuMMI captures complementary material-specific, biochemical, and environmental information. Sample-level quality weights are assigned to ensure full utilization of available data while mitigating low-confidence and sparsely recorded entries. Ablation studies highlight the most influential tabular features, clarifying their contribution to the prediction. Through rigorous external validation across independence-preserving temporal, nanomaterial-held-out, and protein-held-out evaluations, our framework consistently achieves good performance (mean of five classification metrics exceeding 0.75), highlighting its robustness and generalizability to unseen data. Furthermore, fine-tuning on independent gold-nanoparticle data and a held-out protein subset further delivers better performance than training from scratch with substantially fewer samples. Together, our approach enables generalizable and transferable NPI prediction and may accelerate in vitro research and applications of nanomaterials.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。