arXiv:2510.22898cs.AIcs.SE2025-10

提出新框架与测试集,显著提升智能体工具调用的泛化能力

On Generalization in Agentic Tool Calling: CoreThink Agentic Reasoner and MAVEN Dataset

  • 用符号推理层增强大模型,实现任务分解与工具动态调度
  • 在新基准 MAVEN 上多数模型准确率低于50%,暴露泛化短板
  • 无需额外训练,性能提升530%,计算成本仅为十分之一

跨智能体工具调用环境的泛化仍是构建可靠智能体推理系统的关键挑战。尽管大语言模型在孤立基准上表现优异,但其在不同领域间迁移推理策略与协调工具的能力仍不清晰。本文对主流LLM在多个工具调用基准(BFCL v3、TauBench、Tau2Bench、AceBench)进行了大规模评估,并引入MAVEN(Math & Physics Adversarial Verification & Evaluation Network)——一个用于压力测试多步推理的新分布外(OOD)基准,通过显式验证和对抗性任务组合检验模型鲁棒性。结果表明,多数现有模型在MAVEN上的准确率低于50%,揭示了工具使用场景间存在显著泛化差距。为此,我们提出CoreThink智能体推理框架,通过轻量级符号推理层增强LLM,实现结构化任务分解与自适应工具编排。该框架无需额外训练即可在所有基准上泛化,性能相较基线提升530%,计算成本约为十分之一。

原文摘要 · Abstract (English)

Generalization across Agentic tool-calling environments remains a key unsolved challenge in developing reliable agentic reasoning systems. While large language models (LLMs) demonstrate strong performance on isolated benchmarks, their ability to transfer reasoning strategies and co-ordinate tools across diverse domains is poorly understood. In this work, we conduct a large-scale evaluation of state-of-the-art LLMs on multiple tool-calling benchmarksBFCL v3, TauBench, Tau2Bench, and AceBenchand introduce MAVEN (Math & Physics Adversarial Verification & Evaluation Network), a new out of distribution (OOD) benchmark designed to stress-test multi-step reasoning through explicit verification and adversarial task composition. Our results show that most current models achieve below 50% accuracy on MAVEN, revealing a significant generalization gap across tool-use settings. To address this, we present the CoreThink Agentic Reasoner, a framework that augments LLMs with a lightweight symbolic reasoning layer for structured decomposition and adaptive tool orchestration. Without additional training, it generalizes across all benchmarks, achieving state-of-the-art performance with 530% improvements over existing baselines at roughly one-tenth the computational cost.

智能体推理工具调用泛化能力符号推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。