工具对化学语言模型的帮助因任务类型而异,通用题型未必需要工具。
ChemToolAgent: The Impact of Tools on Language Agents for Chemistry Problem Solving
- 基于ChemCrow构建新代理,测试工具在不同化学任务中的表现。
- 在考试类通用问题上,无工具模型反而表现更优。
- 合成预测等专业任务需专用工具,通用推理更依赖知识而非工具。
为提升大语言模型在化学问题求解中的能力,已有研究提出如ChemCrow和Coscientist等基于LLM并集成工具的智能体。然而,这些方法的评估范围有限,难以全面理解工具在多样化化学任务中的实际价值。为此,我们开发了ChemToolAgent——在ChemCrow基础上增强的化学智能体,并对其在专业化化学任务与一般化学问题上的表现进行了全面评估。令人意外的是,ChemToolAgent并未在所有任务中持续优于基础无工具的LLM。通过化学专家的错误分析发现:对于合成预测等专业任务,应引入专用工具;但对于类似考试题目的一般性化学问题,正确推理能力比工具使用更为关键,工具增强并不总是有效。
原文摘要 · Abstract (English)
To enhance large language models (LLMs) for chemistry problem solving, several LLM-based agents augmented with tools have been proposed, such as ChemCrow and Coscientist. However, their evaluations are narrow in scope, leaving a large gap in understanding the benefits of tools across diverse chemistry tasks. To bridge this gap, we develop ChemToolAgent, an enhanced chemistry agent over ChemCrow, and conduct a comprehensive evaluation of its performance on both specialized chemistry tasks and general chemistry questions. Surprisingly, ChemToolAgent does not consistently outperform its base LLMs without tools. Our error analysis with a chemistry expert suggests that: For specialized chemistry tasks, such as synthesis prediction, we should augment agents with specialized tools; however, for general chemistry questions like those in exams, agents' ability to reason correctly with chemistry knowledge matters more, and tool augmentation does not always help.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。