用真实经济竞争测试大模型的商业能力,发现少数模型能赚钱,多数只能保本。
Market-Bench: Benchmarking Large Language Models on Economic and Trade Competition

- 让大模型扮演零售商,在预算受限拍卖中竞购商品
- 只有少数模型能持续盈利,多数虽文案不错仍难逃盈亏平衡
- 适合研究智能体在市场中的博弈行为,尤其关注经济决策
大语言模型管理与获取经济资源的能力尚不明确。本文提出Market-Bench,一个通过经济与贸易竞争来评估大模型在相关任务中表现的综合性基准。我们构建了一个可配置的多智能体供应链经济模型,其中大模型作为零售商代理负责采购与销售商品。在采购阶段,大模型需在预算受限的拍卖中竞标有限库存;在零售阶段,它们设定零售价、生成营销口号,并通过基于角色的关注机制向买家提供信息以促成购买。Market-Bench记录完整的出价、价格、口号、销量及资产负债状态轨迹,支持经济、运营和语义指标的自动评估。对20个开源与闭源大模型代理的基准测试显示显著性能差异,出现赢家通吃的趋势:仅少数模型能持续实现资本增值,而许多模型即使具备相似的语义匹配分数,仍长期处于盈亏平衡点附近。该基准为研究大模型在竞争性市场中的交互行为提供了可复现的测试平台。
原文摘要 · Abstract (English)
The ability of large language models (LLMs) to manage and acquire economic resources remains unclear. In this paper, we introduce \textbf{Market-Bench}, a comprehensive benchmark that evaluates the capabilities of LLMs in economically-relevant tasks through economic and trade competition. Specifically, we construct a configurable multi-agent supply chain economic model where LLMs act as retailer agents responsible for procuring and retailing merchandise. In the \textbf{procurement} stage, LLMs bid for limited inventory in budget-constrained auctions. In the \textbf{retail} stage, LLMs set retail prices, generate marketing slogans, and provide them to buyers through a role-based attention mechanism for purchase. Market-Bench logs complete trajectories of bids, prices, slogans, sales, and balance-sheet states, enabling automatic evaluation with economic, operational, and semantic metrics. Benchmarking on 20 open- and closed-source LLM agents reveals significant performance disparities and winner-take-most phenomenon, \textit{i.e.}, only a small subset of LLM retailers can consistently achieve capital appreciation, while many hover around the break-even point despite similar semantic matching scores. Market-Bench provides a reproducible testbed for studying how LLMs interact in competitive markets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。