arXiv:2605.30738cs.AI2026-05被引 1

用轻量符号框架提升智能体跨领域工具调用的泛化能力

MAVEN: Improving Generalization in Agentic Tool Calling

论文配图:MAVEN: Improving Generalization in Agentic Tool Calling
图 1 · 摘自论文原文
  • 设计模块化验证执行网络,分步分解任务并动态协调工具
  • 在多步数学物理推理任务中,准确率从48%提升至71%
  • 无需额外训练,适合追求高效可靠推理的开发者

智能体在不同工具调用环境中的泛化能力仍是可靠推理系统的核心挑战。尽管大语言模型在单个基准上表现优异,但其组合推理策略、保持中间状态及跨领域工具协调的能力仍待深入探索。我们提出MAVEN(Modular Agentic Verification and Execution Network),一种轻量级符号推理框架,支持结构化分解、自适应工具编排和中间结果验证。我们在多个现有工具调用基准(包括BFCL v3、TauBench、Tau2Bench、AceBench)上评估MAVEN,并引入MAVEN-Bench,一个针对多步数学与物理推理的应力测试基准,包含显式验证和对抗性任务设计。在直接测试中,MAVEN将GPT-OSS-120b基模型的准确率从48%提升至71%,且不需额外训练。相比前沿专有模型,其性能相当,但使用开源权重,成本仅为约1/10,表明轻量级验证中心框架可有效增强组合推理能力,并推动对真实场景中智能体的更过程感知评估。

原文摘要 · Abstract (English)

Generalization across agentic tool-calling environments remains a central challenge for reliable agentic reasoning systems. Although large language models achieve strong results on individual benchmarks, their ability to compose reasoning strategies, preserve intermediate states, and coordinate tools across domains remains underexplored. We present MAVEN (Modular Agentic Verification and Execution Network), a lightweight symbolic reasoning scaffold for structured decomposition, adaptive tool orchestration, and intermediate verification. We evaluate MAVEN across established tool-calling benchmarks, including BFCL v3, TauBench, Tau2Bench, AceBench, and introduce MAVEN-Bench, a stress-test benchmark for multi-step mathematical and physical reasoning with explicit verification and adversarial task composition. MAVEN-Bench exposes a substantial gap between partial reasoning quality and end-to-end task success; in direct MAVEN-Bench runs, MAVEN improves its GPT-OSS-120b base model from 48% to 71% accuracy without additional training. It also remains competitive with frontier proprietary baselines while using an open-weight backbone with an estimated cost ratio of roughly 1/10, suggesting that lightweight verification-centered scaffolds can strengthen compositional reasoning and motivate more process-aware evaluation of agents in the wild.

智能体工具调用推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。