arXiv:2505.05577cs.LGcs.AI2025-05ICML被引 1

PyTDC统一生物多模态模型训练评估,助力药物研发新范式

PyTDC: A multimodal machine learning training, evaluation, and inference platform for biomedical foundation models

  • 构建统一平台整合多源生物数据与多种机器学习任务
  • 首个单细胞药物靶点推荐任务基准,验证模型泛化能力瓶颈
  • 适合生物医学AI研究者开发多模态基础模型

现有生物医学基准缺乏端到端的基础设施,难以支持融合多模态生物数据与广泛机器学习任务的治疗领域模型训练、评估与推理。我们提出PyTDC,一个开源机器学习平台,提供针对多模态生物AI模型的统一训练、评估与推理工具链。PyTDC整合分布式、异构、持续更新的数据源与模型权重,标准化基准测试与推理接口。本文阐述了PyTDC架构组件,并首次开展单细胞药物靶点推荐这一全新机器学习任务的案例研究。结果显示,图表示学习与图论领域专用方法在此任务上表现不佳。尽管上下文感知的几何深度学习方法优于现有最先进(SoTA)及领域基线方法,但其仍无法泛化至未见细胞类型或融合额外模态,凸显出构建多模态、上下文感知的基础模型在生物医学人工智能中的巨大潜力。

原文摘要 · Abstract (English)

Existing biomedical benchmarks do not provide end-to-end infrastructure for training, evaluation, and inference of models that integrate multimodal biological data and a broad range of machine learning tasks in therapeutics. We present PyTDC, an open-source machine-learning platform providing streamlined training, evaluation, and inference software for multimodal biological AI models. PyTDC unifies distributed, heterogeneous, continuously updated data sources and model weights and standardizes benchmarking and inference endpoints. This paper discusses the components of PyTDC's architecture and, to our knowledge, the first-of-its-kind case study on the introduced single-cell drug-target nomination ML task. We find state-of-the-art methods in graph representation learning and domain-specific methods from graph theory perform poorly on this task. Though we find a context-aware geometric deep learning method that outperforms the evaluated SoTA and domain-specific baseline methods, the model is unable to generalize to unseen cell types or incorporate additional modalities, highlighting PyTDC's capacity to facilitate an exciting avenue of research developing multimodal, context-aware, foundation models for open problems in biomedical AI.

生物医疗AI多模态学习药物发现深度学习平台

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。