构建科学工具使用评测基准,提升大模型跨领域科研协作能力
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents
- 设计可扩展交互环境,集成1780个跨学科科学工具
- 发现主流模型在长程任务中性能显著下降
- 提出依赖图驱动的数据生成法,实现小模型超越大模型
科学推理本质上需要整合复杂的工具集来处理领域知识。然而,现有评测基准大多忽视了智能体协调工具以完成严谨工作流的能力。为此,我们提出了SciAgentGym,一个可扩展的交互环境,涵盖四个自然科学领域中的1,780个领域专用工具,并配有稳健的执行基础设施。同时,我们设计了分层评估套件SciAgentBench,用于从基础操作到长程工作流全面测试智能体能力。评估发现:当前最先进模型在复杂科学工具使用上仍存在严重瓶颈,其性能随交互时长增加而大幅下降。为解决此问题,我们提出SciForge,一种将工具动作空间建模为依赖图的数据合成方法,生成逻辑感知的训练轨迹。通过在这些轨迹上微调,我们的SciAgent-8B在性能上超越显著更大的Qwen3-VL-235B-Instruct,并展现出良好的跨领域迁移能力。结果表明下一代自主科学智能体具有巨大潜力。
原文摘要 · Abstract (English)
Scientific reasoning inherently demands integrating sophisticated toolkits to navigate domain-specific knowledge. Yet, current benchmarks largely overlook agents' ability to orchestrate tools for such rigorous workflows. To bridge this gap, we introduce SciAgentGym, a scalable interactive environment featuring 1,780 domain-specific tools across four natural science disciplines, supported by a robust execution infrastructure. Complementing this, we present SciAgentBench, a tiered evaluation suite designed to stress-test agentic capabilities from elementary actions to long-horizon workflows. Our evaluation identifies a critical bottleneck: state-of-the-art models still struggle with complex scientific tool-use, and their performance degrades substantially as interaction horizons extend. To address this, we propose SciForge, a data synthesis method that models the tool action space as a dependency graph to generate logic-aware training trajectories. By fine-tuning on these trajectories, our SciAgent-8B outperforms the significantly larger Qwen3-VL-235B-Instruct while exhibiting positive cross-domain transfer of scientific tool-use capabilities. These results underscore the promising potential of next-generation autonomous scientific agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。