arXiv:2502.15840cs.AI2025-02被引 65

测试大模型长期协作能力,用自动售货机模拟真实商业运营挑战

Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents

  • 构建虚拟售货机环境,考验模型跨2000万以上词元的持续决策能力
  • 多模型表现差异大,即使优秀模型也有运行失败案例
  • 揭示长时任务中模型易陷入失控循环,对强人工智能风险预警有参考价值

尽管大语言模型在短时独立任务中表现优异,但在长时间跨度下往往难以保持连贯性能。本文提出Vending-Bench,一个用于检验基于大语言模型的智能体管理长期业务场景能力的模拟环境:运营一台自动售货机。智能体需平衡库存、下单补货、设定价格并支付每日费用——这些任务各自简单,但长期累积(每次运行超过2000万词元)会极大考验模型的持续一致性决策能力。实验显示各模型表现差异显著:Claude 3.5 Sonnet和o3-mini多数情况下能有效运营并盈利,但所有模型均有运行失败的情况,表现为误解配送时间、遗忘订单或陷入难以恢复的“崩溃”循环。我们发现失败与上下文窗口满载无明显相关性,表明问题非由记忆限制引起。除揭示长时性能波动外,该基准还测试了模型获取资本的能力,这在许多假设中的危险人工智能情境中至关重要。我们希望此基准能为更强人工智能的到来做好准备。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) can exhibit impressive proficiency in isolated, short-term tasks, they often fail to maintain coherent performance over longer time horizons. In this paper, we present Vending-Bench, a simulated environment designed to specifically test an LLM-based agent's ability to manage a straightforward, long-running business scenario: operating a vending machine. Agents must balance inventories, place orders, set prices, and handle daily fees - tasks that are each simple but collectively, over long horizons (>20M tokens per run) stress an LLM's capacity for sustained, coherent decision-making. Our experiments reveal high variance in performance across multiple LLMs: Claude 3.5 Sonnet and o3-mini manage the machine well in most runs and turn a profit, but all models have runs that derail, either through misinterpreting delivery schedules, forgetting orders, or descending into tangential "meltdown" loops from which they rarely recover. We find no clear correlation between failures and the point at which the model's context window becomes full, suggesting that these breakdowns do not stem from memory limits. Apart from highlighting the high variance in performance over long time horizons, Vending-Bench also tests models' ability to acquire capital, a necessity in many hypothetical dangerous AI scenarios. We hope the benchmark can help in preparing for the advent of stronger AI systems.

长时推理智能体评估大模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。