用动态商战模拟测试大模型长期决策能力,发现其表现差异显著。
AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic Simulations
- 设计月度滚动的零售企业模拟器,让模型做12个月战略决策。
- 模型在利润、营收和市场份额上差距明显,最佳仅达人类水平80%。
- 适合研究AI管理能力或评估模型长期规划能力的学者使用。
大语言模型(LLMs)在自然语言处理方面表现优异,但在多步、战略性的商业决策领域仍缺乏系统评估。本文提出一个可复现、开源的管理模拟器,用于评估五款主流免费在线大模型(Gemini、ChatGPT、Meta AI、Mistral AI、Grok)在动态零售企业模拟中的表现。每期提供完整业务报告,模型需制定定价、采购量、营销预算、招聘、裁员、贷款、培训、研发、销售与收入预测等决策。通过12个月的月度滚动模拟,以利润、营收、市场份额等量化指标对比模型表现,并分析其策略一致性、适应性及决策理由。结果表明,模型在长期决策中存在显著差异,最优模型仅达到人类基准的80%。该框架为评估大模型在复杂商业环境下的长期规划能力提供了新范式。
原文摘要 · Abstract (English)
The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent trends in AI benchmarking is performance of Large Language Models (LLMs) over longer time horizons. While LLMs excel at tasks involving natural language and pattern recognition, their capabilities in multi-step, strategic business decision-making remain largely unexplored. Few studies demonstrated how results can be different from benchmarks in short-term tasks, as Vending-Bench revealed. Meanwhile, there is a shortage of alternative benchmarks for long-term coherence. This research analyses a novel benchmark using a business game for the decision making in business. The research contributes to the recent literature on AI by proposing a reproducible, open-access management simulator to the research community for LLM benchmarking. This novel framework is used for evaluating the performance of five leading LLMs available in free online interface: Gemini, ChatGPT, Meta AI, Mistral AI, and Grok. LLM makes decisions for a simulated retail company. A dynamic, month-by-month management simulation provides transparently in spreadsheet model as experimental environment. In each of twelve months, the LLMs are provided with a structured prompt containing a full business report from the previous period and are tasked with making key strategic decisions: pricing, order size, marketing budget, hiring, dismissal, loans, training expense, R&D expense, sales forecast, income forecast The methodology is designed to compare the LLMs on quantitative metrics: profit, revenue, and market share, and other KPIs. LLM decisions are analyzed in their strategic coherence, adaptability to market changes, and the rationale provided for their decisions. This approach allows to move beyond simple performance metrics for assessment of the long-term decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。