探究大模型智能体在药物发现中的模块化设计,发现组件替换需谨慎。
Exploring Modularity of Agentic Systems for Drug Discovery
- 对比不同大模型与智能体类型,评估其可互换性。
- Claude-3.5-Sonnet等模型表现优于Llama、GPT-3.5等。
- 代码生成智能体整体更优,但效果依赖具体问题和模型。
大语言模型(LLMs)与智能体系统为加速药物发现带来新机遇。本文研究基于大模型的智能体系统在药物发现中的模块化问题,即模型与智能体类型是否可互换,这一问题在药物发现领域关注较少。我们对比了不同大模型及工具调用智能体与代码生成智能体的性能。案例研究通过LLM-as-a-judge评分,评估其在化学与药物发现中协调工具的能力,结果显示Claude-3.5-Sonnet、Claude-3.7-Sonnet和GPT-4o优于Llama-3.1-8B、Llama-3.1-70B、GPT-3.5-Turbo和Nova-Micro。尽管代码生成智能体平均表现更优,但其效果高度依赖问题类型与模型。此外,更换系统提示的影响也因问题和模型而异,表明即使在特定领域,也不能随意替换系统组件而不重新设计。本研究强调需进一步探索智能体系统的模块化,以推动真实世界应用的可靠与模块化解决方案发展。
原文摘要 · Abstract (English)
Large-language models (LLMs) and agentic systems present exciting opportunities to accelerate drug discovery. In this study, we examine the modularity of LLM-based agentic systems for drug discovery, i.e., whether parts of the system such as the LLM and type of agent are interchangeable, a topic that has received limited attention in drug discovery. We compare the performance of different LLMs and the effectiveness of tool-calling agents versus code-generating agents. Our case study, comparing performance in orchestrating tools for chemistry and drug discovery using an LLM-as-a-judge score, shows that Claude-3.5-Sonnet, Claude-3.7-Sonnet and GPT-4o outperform alternative language models such as Llama-3.1-8B, Llama-3.1-70B, GPT-3.5-Turbo, and Nova-Micro. Although we confirm that code-generating agents outperform the tool-calling ones on average, we show that this is highly question- and model-dependent. Furthermore, the impact of replacing system prompts is dependent on the question and model, underscoring that even in this particular domain one cannot just replace components of the system without re-engineering. Our study highlights the necessity of further research into the modularity of agentic systems to enable the development of reliable and modular solutions for real-world problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。