arXiv:2508.10899cs.LG2025-08NeurIPS被引 2

从文献中提取药物设计先验知识,提升AI生成分子的安全性与有效性。

A Dataset for Distilling Knowledge Priors from Literature for Therapeutic Design

  • 用大模型从文献中挖掘治疗分子的自然语言事实与结构表示
  • 3230万条数据对使小模型性能超越大模型,分类回归任务均优于基线
  • 可作为约束条件优化新分子设计,显著降低致突变风险

AI驱动的新药发现能大幅缩短设计周期并提升疗效。使用模拟器的模型虽可探索广阔设计空间,但因缺乏实验先验,常违反隐含约束。我们在GuacaMol基准上对多种模型进行新分析发现,超过60%提出的分子具有高致突变概率。本文提出Medex数据集,通过大模型管道从实验室相关文献中提取治疗实体信息,并总结为简洁的合理使用事实。Medex包含3230万对自然语言事实与对应实体表示(如SMILES或RefSeq ID)。我们训练了LLM、CLIP和LLava架构,在Therapeutic Data Commons(TDC)任务上评估其联合文本与设计目标的推理能力。以该数据预训练的小模型(1500万参数)在回归与分类任务上表现优于20亿参数的TxGemma,平均接近90亿参数模型。基于Medex构建的模型可用于GuacaMol中新分子优化时施加安全约束,生成更安全且几乎同样有效的分子。数据集已发布于https://huggingface.co/datasets/medexanon/Medex,随文献增长将持续扩展。

原文摘要 · Abstract (English)

AI-driven discovery can greatly reduce design time and enhance new therapeutics' effectiveness. Models using simulators explore broad design spaces but risk violating implicit constraints due to a lack of experimental priors. For example, in a new analysis we performed on a diverse set of models on the GuacaMol benchmark using supervised classifiers, over 60\% of molecules proposed had high probability of being mutagenic. In this work, we introduce Medex, a dataset of priors for design problems extracted from literature describing compounds used in lab settings. It is constructed with LLM pipelines for discovering therapeutic entities in relevant paragraphs and summarizing information in concise fair-use facts. Medex consists of 32.3 million pairs of natural language facts, and appropriate entity representations (i.e. SMILES or refseq IDs). To demonstrate the potential of the data, we train LLM, CLIP, and LLava architectures to reason jointly about text and design targets and evaluate on tasks from the Therapeutic Data Commons (TDC). Medex is highly effective for creating models with strong priors: in supervised prediction problems that use our data as pretraining, our best models with 15M learnable parameters outperform larger 2B TxGemma on both regression and classification TDC tasks, and perform comparably to 9B models on average. Models built with Medex can be used as constraints while optimizing for novel molecules in GuacaMol, resulting in proposals that are safer and nearly as effective. We release our dataset at https://huggingface.co/datasets/medexanon/Medex, and will provide expanded versions as available literature grows.

药物发现知识蒸馏大模型应用分子生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。