用拍卖机制动态分配任务,提升大模型代理的推理效率。
Agora: Enhancing LLM Agent Reasoning Via Auction-Based Task Allocation

- 将推理步骤视为可交易物品,基于校准后能力分配任务。
- 在五个基准上优于或媲美单模型、路由和级联基线。
- 适合需要高效调用多模型工具的复杂推理场景。
提升大型语言模型(LLM)代理的推理能力,需有效协调多种专家模型与工具。然而,现有框架通常基于粗粒度的任务-功能匹配调用API,忽视了功能相似替代方案之间的性能波动与成本效率差异。为此,我们提出Agora,一个基于置信度校准拍卖的动态任务分配框架。通过将推理步骤视为可交易物品,Agora依据校准后的胜任力而非原始置信度进行分配。在五个主要基准测试中,当候选池匹配时,Agora的表现优于或媲美单模型、路由与级联基线。
原文摘要 · Abstract (English)
Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse expert models and tools. However, existing frameworks typically call APIs, based on coarse-grained matching between tasks and the functions of expert models or tools, while overlooking critical factors such as performance variability and cost efficiency among functionally similar alternatives. To address this, we propose Agora, a framework that uses a confidence-calibrated auction to dynamically allocate tasks to expert models and tools. By treating reasoning steps as tradeable items, Agora bases allocation on calibrated competence rather than raw confidence. Across five main benchmarks, Agora improves or remains competitive with single-model, routing, and cascade baselines under matched candidate pools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。