构建首个真实药物分子设计多轮评估基准,检验大模型自主研发能力
SMDD-Bench: Can LLMs Solve Real-World Small Molecule Drug Design Tasks?

- 设计502个跨化学空间的真实药物任务,需多轮推理与规划
- 最强模型仅解决40.2%任务,暴露当前大模型在药物设计中的局限性
- 适合关注药物研发自动化、大模型科学推理的科研人员参考
大语言模型代理在科学发现中潜力巨大,但其在真实世界小分子药物设计任务上的表现尚不清晰。现有评估方法或过于简单、或规模有限、或仅限单轮问答。为标准化评估,我们提出SMDD-Bench,一个具有挑战性的多轮、长周期代理基准,包含502个可解任务实例,涵盖2D药效团识别、相互作用点发现、骨架跃迁、先导优化和片段组装五类任务,覆盖102个独特蛋白靶点,横跨广泛化学空间。完全解决该基准需具备强化学与生物学推理能力、三维直觉、专业工具使用及有限调用下规划能力。我们评测了7个前沿开源与闭源大模型,发现即使最强模型GPT5.4也仅能解决40.2%的任务。我们已在smddbench.com发布公开排行榜,旨在推动实现全自动计算药物设计。
原文摘要 · Abstract (English)
LLM agents have incredible potential for scientific discovery applications. However, the performance of LLM agents on real-world, small molecule drug design (SMDD) tasks across diverse chemistries and targets is unclear. Current evaluation methods are either ad hoc, too simple for real-world discovery, limited in scale, or restricted to single-turn question answering. In effort to standardize the evaluation of LLM agents on small molecule design, we introduce SMDD-Bench, a challenging, multi-turn, long-horizon agentic benchmark consisting of 502 guaranteed-solvable task instances spanning 5 task types: 2D Pharmacophore Identification, Interaction Point Discovery, Scaffold Hopping, Lead Optimization, and Fragment Assembly. SMDD-Bench tasks span a wide region of chemical space and involve 102 unique protein targets. Completely solving the benchmark would require having strong chemical and biological reasoning and 3D intuition, understanding specialized tool use, and displaying planning expertise over a limited number of oracle calls. We benchmark 7 frontier open and closed source LLMs and find even the most performant LLM, GPT5.4, solves only 40.2\% of tasks. We hope SMDD-Bench provides a standardized testbed to invigorate the field towards training and evaluating LLM agents for fully autonomous computational drug design. We host a public leaderboard at smddbench.com .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。