让不同模型协作如团队,提升推理与编程准确率。
Team of Thoughts: Efficient Test-time Scaling of Agentic Systems through Orchestrated Tool Calling
- 用调度器管理多种模型,按专长动态组合。
- 数学与代码任务准确率达96%和77.91%,超越单一模型。
- 适合需要多模型协同的复杂任务研究者。
现有多智能体系统(MAS)通常采用同质模型配置,未能利用不同后训练架构的多样专长。我们提出 Team-of-Thoughts,一种异构 MAS 框架,将不同模型视为调度器驱动下的专用工具。该框架引入两个新组件:(1) 调度器校准,识别具备优秀协调与综合能力的模型;(2) 代理自评估协议,让工具代理自主标注自身领域优势以指导选择。推理时,调度器根据这些画像动态激活最适配的代理,以最大化能力覆盖。在五个数学推理与代码生成基准上,Team-of-Thoughts 持续优于单个模型和现有 MAS 基线。特别地,在 AIME24 与 LiveCodeBench 上,分别达到 96.00% 和 77.91% 的准确率,显著优于同质角色扮演基线(80.00% 和 65.93%)。
原文摘要 · Abstract (English)
Existing Multi-Agent Systems (MAS) typically rely on homogeneous model configurations, failing to exploit the diverse expertise inherent in different post-trained architectures. We propose Team-of-Thoughts, a heterogeneous MAS framework that treats diverse models as specialized tools within an orchestrator-driven paradigm. Team-of-Thoughts introduces two novel components: (1) Orchestrator Calibration, which identifies models with superior coordination and synthesis capabilities, and (2) Agent Self-Assessment, a protocol where tool agents profile their own domain-specific strengths to guide selection. At inference, the orchestrator dynamically activates the most compatible agents based on these profiles to maximize capability coverage. Across five mathematical reasoning and code generation benchmarks, Team-of-Thoughts consistently outperforms individual models and existing MAS baselines. Notably, on AIME24 and LiveCodeBench, Team-of-Thoughts achieves 96.00% and 77.91% accuracy, respectively, significantly improving over homogeneous role-play baselines (80.00% and 65.93%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。