arXiv:2602.13209q-fin.GNcs.AI2026-02

用柠檬水摊模拟小生意,测试大模型的经济直觉与决策能力。

LemonadeBench: Evaluating the Economic Intuition of Large Language Models in Simple Markets

  • 通过模拟30天柠檬水摊运营,评估大模型在库存、定价、营业时间上的综合决策能力。
  • 前沿模型利润达理论最优的70%,比基础模型提升超10倍,但仍有明显盲区。
  • 发现模型多局部优化而非全局最优,适合研究经济推理与认知局限的学者。

我们提出LemonadeBench v0.5,一个最小化基准,用于通过模拟柠檬水摊业务,评估大语言模型(LLMs)在简单市场中的经济直觉、长期规划及不确定性下的决策能力。模型需管理具有保质期的商品库存,设定价格,决定营业时间,并在30天内最大化利润——这些任务是小企业主日常面临的挑战。所有模型均表现出有意义的经济自主性,实现盈利;性能随模型复杂度显著提升,从基础模型获得微薄利润到前沿模型捕捉理论最优利润的70%,改善超过10倍。然而,对六维度商业效率的分解显示一致模式:模型仅在部分领域表现优异,存在显著盲点,体现局部而非全局优化特征。

原文摘要 · Abstract (English)

We introduce LemonadeBench v0.5, a minimal benchmark for evaluating economic intuition, long-term planning, and decision-making under uncertainty in large language models (LLMs) through a simulated lemonade stand business. Models must manage inventory with expiring goods, set prices, choose operating hours, and maximize profit over a 30-day period-tasks that any small business owner faces daily. All models demonstrate meaningful economic agency by achieving profitability, with performance scaling dramatically by sophistication-from basic models earning minimal profits to frontier models capturing 70% of theoretical optimal, a greater than 10x improvement. Yet our decomposition of business efficiency across six dimensions reveals a consistent pattern: models achieve local rather than global optimization, excelling in select areas while exhibiting surprising blind spots elsewhere.

经济推理大模型评估决策能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。