构建58个微观经济推理要素的评估体系,测试大模型在供需分析等非策略场景中的表现。
STEER-ME: Assessing the Microeconomic Reasoning of Large Language Models
- 将微观经济推理拆解为58个要素,覆盖10个领域、5种视角和3类逻辑
- 提出auto-STEER生成协议,自动构造新题避免模型过拟合评测集
- 测试27个大模型在多种提示策略下的表现,提供可复用的评估框架
如何判断大语言模型能否可靠进行经济推理?现有大多数基准测试聚焦特定应用场景,未能涵盖丰富的经济任务。尽管Raman等人[2024]提出了全面评估策略决策的方法,但未涉及微观经济学中常见的非策略场景,如供需分析。本文通过将微观经济推理细分为58个独立要素,聚焦供需逻辑,每个要素基于最多10个不同领域、5种视角和3种类型。通过新型的LLM辅助数据生成协议auto-STEER,实现跨组合空间的自动化题目生成,该方法通过适配手写模板以适应新领域和视角。由于能持续生成新问题,auto-STEER有效降低模型对评测集过拟合的风险。我们通过27个大模型(从开源小模型到当前最先进模型)的案例研究验证了该基准的实用性,在不同提示策略与评分指标下评估模型解决微观经济问题的能力。
原文摘要 · Abstract (English)
How should one judge whether a given large language model (LLM) can reliably perform economic reasoning? Most existing LLM benchmarks focus on specific applications and fail to present the model with a rich variety of economic tasks. A notable exception is Raman et al. [2024], who offer an approach for comprehensively benchmarking strategic decision-making; however, this approach fails to address the non-strategic settings prevalent in microeconomics, such as supply-and-demand analysis. We address this gap by taxonomizing microeconomic reasoning into $58$ distinct elements, focusing on the logic of supply and demand, each grounded in up to $10$ distinct domains, $5$ perspectives, and $3$ types. The generation of benchmark data across this combinatorial space is powered by a novel LLM-assisted data generation protocol that we dub auto-STEER, which generates a set of questions by adapting handwritten templates to target new domains and perspectives. Because it offers an automated way of generating fresh questions, auto-STEER mitigates the risk that LLMs will be trained to over-fit evaluation benchmarks; we thus hope that it will serve as a useful tool both for evaluating and fine-tuning models for years to come. We demonstrate the usefulness of our benchmark via a case study on $27$ LLMs, ranging from small open-source models to the current state of the art. We examined each model's ability to solve microeconomic problems across our whole taxonomy and present the results across a range of prompting strategies and scoring metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。