arXiv:2410.17950cs.AI2024-10被引 1

用更小的模型实现更强函数调用能力,提升效率与可靠性。

Benchmarking Floworks against OpenAI & Anthropic: A Novel Framework for Enhanced LLM Function Calling

  • 设计ThorV2架构,优化大模型函数调用逻辑
  • 在HubSpot CRM任务中准确率、延迟和成本全面优于OpenAI与Anthropic模型
  • 适合构建高效智能助手,尤其擅长多步骤复杂任务

大语言模型(LLMs)在多个领域展现强大能力,但其经济价值受限于工具使用与函数调用难题。本文提出ThorV2新架构,显著提升LLMs的函数调用性能。我们构建了聚焦HubSpot CRM操作的综合基准,评估ThorV2在OpenAI与Anthropic领先模型中的表现。结果表明,无论单次还是多API调用任务,ThorV2在准确性、可靠性、延迟及成本效率方面均优于现有模型。此外,相比传统模型,ThorV2在多步任务中表现更稳定且可扩展性更强。本研究揭示:采用更小规模的LLM,即可实现优于当前最优模型的函数调用精度,对开发更强大智能助手及推动大模型在真实场景应用具有重要意义。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown remarkable capabilities in various domains, yet their economic impact has been limited by challenges in tool use and function calling. This paper introduces ThorV2, a novel architecture that significantly enhances LLMs' function calling abilities. We develop a comprehensive benchmark focused on HubSpot CRM operations to evaluate ThorV2 against leading models from OpenAI and Anthropic. Our results demonstrate that ThorV2 outperforms existing models in accuracy, reliability, latency, and cost efficiency for both single and multi-API calling tasks. We also show that ThorV2 is far more reliable and scales better to multistep tasks compared to traditional models. Our work offers the tantalizing possibility of more accurate function-calling compared to today's best-performing models using significantly smaller LLMs. These advancements have significant implications for the development of more capable AI assistants and the broader application of LLMs in real-world scenarios.

函数调用大模型优化智能助手性能评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。