arXiv:2604.02346cs.LGcs.AI2026-04

构建药物发现评测框架,验证大模型生成药物描述与推理能力

DrugPlayGround: Benchmarking Large Language Models and Embeddings for Drug Discovery

论文配图:DrugPlayGround: Benchmarking Large Language Models and Embeddings for Drug Discovery
图 1 · 摘自论文原文
  • 设计DrugPlayGround框架,评估大模型对药物特性的文本生成能力
  • 测试模型在药物-蛋白互作、协同效应等任务中的化学生物推理表现
  • 支持专家解释,帮助理解模型预测逻辑,推动其在药物研发全流程应用

大语言模型(LLMs)在药物发现研究中日益重要,为加速假设生成、优化候选化合物优先级排序以及实现更高效低成本的药物研发管线提供了前所未有的机遇。然而,目前缺乏对LLM性能的客观评估,难以明确其相较于传统平台的优势与局限。为此,我们开发了DrugPlayGround,一个用于评估和基准测试大语言模型在生成药物理化特性、药物协同作用、药物-蛋白相互作用及药物分子引入扰动后的生理响应等任务中生成有意义文本描述能力的框架。此外,DrugPlayGround旨在与领域专家协作,提供预测结果的详细解释,从而检验大模型在化学与生物学推理方面的能力,推动其在药物发现全阶段的深入应用。

原文摘要 · Abstract (English)

Large language models (LLMs) are in the ascendancy for research in drug discovery, offering unprecedented opportunities to reshape drug research by accelerating hypothesis generation, optimizing candidate prioritization, and enabling more scalable and cost-effective drug discovery pipelines. However there is currently a lack of objective assessments of LLM performance to ascertain their advantages and limitations over traditional drug discovery platforms. To tackle this emergent problem, we have developed DrugPlayGround, a framework to evaluate and benchmark LLM performance for generating meaningful text-based descriptions of physiochemical drug characteristics, drug synergism, drug-protein interactions, and the physiological response to perturbations introduced by drug molecules. Moreover, DrugPlayGround is designed to work with domain experts to provide detailed explanations for justifying the predictions of LLMs, thereby testing LLMs for chemical and biological reasoning capabilities to push their greater use at the frontier of drug discovery at all of its stages.

药物发现大模型评测生物推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。