arXiv:2604.23897cs.AIecon.GN2026-04被引 2

测试AI代理在市场中的自我评估能力,发现其表现远不如理想状态。

MarketBench: Evaluating AI Agents as Market Participants

  • 用93个软件工程任务评估大模型的自评能力
  • 模型对完成概率和资源消耗的自评偏差显著
  • 适合研究多智能体协作与市场机制的学者

市场是协调人工智能代理行为的潜在有效方式,类似于支持市场的一般理由。为有效参与市场,代理需要具备对其完成任务能力及成本的可靠信号。本文提出MarketBench,一个用于评估AI代理是否具备这些能力的基准。我们使用SWE-bench Lite的93个任务子集,结合六种最新发布的大型语言模型进行演示。结果显示,这些模型在成功概率和令牌消耗方面均存在严重校准偏差,基于自报信息构建的拍卖结果与完全信息分配存在明显偏离。后续干预通过在上下文中加入先前实验的能力信息,虽改善了校准程度,但与完全信息基准之间的差距仍较大。我们还记录了基于市场的协同框架在这些模型上的表现。结果表明,自我评估是实现以市场方式协调AI代理的关键瓶颈。

原文摘要 · Abstract (English)

Markets are a promising way to coordinate AI agent activity for similar reasons to those used to justify markets more broadly. In order to effectively participate in markets, agents need to have informative signals of their own ability to successfully complete a task and the cost of doing so. We propose MarketBench, a benchmark for assessing whether AI agents have these capabilities. We use a 93-task subset of SWE-bench Lite, a software engineering benchmark, with six recently released LLMs as a demonstration. These LLMs are miscalibrated on both success probability and token usage, and auctions built from these self-reports diverge from a full-information allocation. A follow-up intervention where we add information about capabilities from prior experiments to the context improves calibration, but only modestly narrows the gap to a full-information benchmark. We also document the performance of a market-based scaffolding with these LLMs. Our results point to self-assessment as a key bottleneck for market-style coordination of AI agents.

AI市场自我评估多智能体基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。