arXiv:2604.00137cs.AIcs.SE2026-04被引 1

构建开源工具社区,提升智能体使用工具的可靠性与安全性。

Open, Reliable, and Collective: A Community-Driven Framework for Tool-Using AI Agents

  • 建立标准化工具接口,支持社区协作维护与评估。
  • 任务特定工具使智能体性能提升6%至22%。
  • 适合关注工具可信性与开源协作的研究者与开发者。

集成工具的大型语言模型可检索信息、执行计算并完成现实任务,但其可靠性不仅依赖于工具调用的准确性,还取决于工具本身的质量,包括正确性、稳定性与安全性。以往研究多关注工具使用,而忽视了工具内在质量。本文提出 OpenTools,一个由社区驱动且可持续维护的开源工具箱,支持发现、使用、评估和贡献开源工具。OpenTools 标准化工具接口,将文档化的 Python 函数转化为可审查的工具包,支持维护者触发评估,并结合非执行风险检测与可选的 LLM 审查建议。公开网页演示允许用户运行工具与智能体、查看证据、提交测试、提交工具供维护者审核,同时 MCP 实现外部应用的受控访问。实验表明,在多种智能体架构下,社区贡献的任务专用工具相比现有工具箱性能提升 6% 至 22%,凸显了工具内在准确性的重要性。

原文摘要 · Abstract (English)

Tool-integrated LLMs retrieve information, perform computations, and take real-world actions, but their reliability depends on both tool-use accuracy and intrinsic tool accuracy, including tool correctness, stability, and safety. While prior work primarily emphasizes tool use, intrinsic tool accuracy remains underexamined. We introduce OpenTools, a community-driven and maintainable toolbox for discovering, using, evaluating, and contributing open-source tools. OpenTools standardizes tool interfaces, converts documented Python functions into reviewable bundles, supports maintainer-triggered evaluation, and combines non-executing risk inspection with optional advisory LLM review. A public web demo allows users to run tools and agents, inspect evidence, contribute tests, and submit tools for maintainer review, while MCP enables controlled access from external applications. Experiments show that community-contributed, task-specific tools yield relative gains of 6% to 22% over an existing toolbox across multiple agent architectures, highlighting the importance of intrinsic tool accuracy.

智能体开源工具可靠性社区协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。