构建首个多模态药物相互作用数据集,助力精准医疗安全
MUDI: A Multimodal Biomedical Dataset for Understanding Pharmacodynamic Drug-Drug Interactions
- 融合文本、化学式、分子结构与图像,全面表征药物对
- 覆盖31万+药对,含协同、拮抗等三类标注,测试集含未见药对
- 开源数据与基准模型,适合医药AI研究者使用
理解不同药物间的相互作用(药物-药物相互作用,DDI)对于保障患者安全和优化治疗效果至关重要。现有DDI数据集主要依赖文本信息,忽略了反映复杂药物机制的多模态数据。本文提出MUDI——一个大规模多模态生物医学数据集,用于理解药效学药物-药物相互作用。MUDI通过整合药理文本、化学式、分子结构图及图像,对310,532个标注为协同、拮抗或新效应的药物对进行综合表征。关键在于,测试集中包含未在训练中出现的药物对,以有效评估机器学习模型的泛化能力。我们采用晚期融合投票与中期融合策略对基准模型进行了评估。所有数据、标注、评估脚本及基线均已开放获取。
原文摘要 · Abstract (English)
Understanding the interaction between different drugs (drug-drug interaction or DDI) is critical for ensuring patient safety and optimizing therapeutic outcomes. Existing DDI datasets primarily focus on textual information, overlooking multimodal data that reflect complex drug mechanisms. In this paper, we (1) introduce MUDI, a large-scale Multimodal biomedical dataset for Understanding pharmacodynamic Drug-drug Interactions, and (2) benchmark learning methods to study it. In brief, MUDI provides a comprehensive multimodal representation of drugs by combining pharmacological text, chemical formulas, molecular structure graphs, and images across 310,532 annotated drug pairs labeled as Synergism, Antagonism, or New Effect. Crucially, to effectively evaluate machine-learning based generalization, MUDI consists of unseen drug pairs in the test set. We evaluate benchmark models using both late fusion voting and intermediate fusion strategies. All data, annotations, evaluation scripts, and baselines are released under an open research license.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。