AI代理能自主设计药物研发流程,但仍有不稳定问题。
Can AI Agents Design and Implement Drug Discovery Pipelines?
- 构建虚拟筛选基准DO Challenge,测试AI独立设计药物流程能力
- 多代理系统表现优于多数人类团队,但未达专家水平
- 适合关注AI在药物研发中应用前景的研究者
人工智能,特别是基于大语言模型的自主代理系统,为加速药物发现提供了新机遇,可提升计算机模拟效率并减少对昂贵实验的依赖。本文提出DO Challenge基准,用于评估AI代理在复杂虚拟筛选场景下的决策能力。该任务要求系统自主开发、实现并执行高效策略,从大规模数据集中识别有潜力的分子结构,同时需应对化学空间探索、模型选择与资源限制等多重目标挑战。我们还分析了基于该基准的DO Challenge 2025竞赛,展示了人类参赛者的多样化策略。此外,我们提出了Deep Thought多代理系统,在基准测试中表现优异,超越多数人类团队。在所测试的语言模型中,Claude 3.7 Sonnet、Gemini 2.5 Pro和o3在主代理角色中表现最佳,GPT-4o和Gemini 2.0 Flash则在辅助角色中表现出色。尽管前景可观,当前系统性能仍低于专家设计方案,且存在较高不稳定性,凸显了AI驱动方法在药物发现与科研中的潜力与局限。
原文摘要 · Abstract (English)
The rapid advancement of artificial intelligence, particularly autonomous agentic systems based on Large Language Models (LLMs), presents new opportunities to accelerate drug discovery by improving in-silico modeling and reducing dependence on costly experimental trials. Current AI agent-based systems demonstrate proficiency in solving programming challenges and conducting research, indicating an emerging potential to develop software capable of addressing complex problems such as pharmaceutical design and drug discovery. This paper introduces DO Challenge, a benchmark designed to evaluate the decision-making abilities of AI agents in a single, complex problem resembling virtual screening scenarios. The benchmark challenges systems to independently develop, implement, and execute efficient strategies for identifying promising molecular structures from extensive datasets, while navigating chemical space, selecting models, and managing limited resources in a multi-objective context. We also discuss insights from the DO Challenge 2025, a competition based on the proposed benchmark, which showcased diverse strategies explored by human participants. Furthermore, we present the Deep Thought multi-agent system, which demonstrated strong performance on the benchmark, outperforming most human teams. Among the language models tested, Claude 3.7 Sonnet, Gemini 2.5 Pro and o3 performed best in primary agent roles, and GPT-4o, Gemini 2.0 Flash were effective in auxiliary roles. While promising, the system's performance still fell short of expert-designed solutions and showed high instability, highlighting both the potential and current limitations of AI-driven methodologies in transforming drug discovery and broader scientific research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。