为大模型经济决策能力设计评估基准与测试框架。
EconEvals: Benchmarks and Litmus Tests for Economic Decision-Making by LLM Agents
- 基于采购、调度、定价构建环境学习能力测试基准。
- 通过多目标冲突任务量化模型权衡、可靠性和能力得分。
- 可评估大模型在经济场景中的行为逻辑与演进趋势。
我们开发了评估大语言模型(LLMs)经济决策能力与倾向性的方法。首先,基于经济学中的关键问题——采购、调度和定价,构建了测试环境学习能力的基准。其次,提出“试纸测试”框架,通过具有多重冲突目标的简化决策任务,量化大模型的选择行为。每个试纸测试输出三个指标:权衡响应的试纸分数、选择行为一致性的可靠性分数,以及在单一明确目标下的能力分数。对一系列前沿大模型的评估表明:(1)模型能力与倾向随时间演变;(2)从模型选择行为与思维链中获得有意义的经济洞察;(3)验证了试纸测试框架的自洽性、鲁棒性与泛化能力。本工作为大模型在经济决策中更深度集成提供了评估基础。
原文摘要 · Abstract (English)
We develop evaluation methods for measuring the economic decision-making capabilities and tendencies of LLMs. First, we develop benchmarks derived from key problems in economics -- procurement, scheduling, and pricing -- that test an LLM's ability to learn from the environment in context. Second, we develop the framework of litmus tests, evaluations that quantify an LLM's choice behavior on a stylized decision-making task with multiple conflicting objectives. Each litmus test outputs a litmus score, which quantifies an LLM's tradeoff response, a reliability score, which measures the coherence of an LLM's choice behavior, and a competency score, which measures an LLM's capability at the same task when the conflicting objectives are replaced by a single, well-specified objective. Evaluating a broad array of frontier LLMs, we (1) investigate changes in LLM capabilities and tendencies over time, (2) derive economically meaningful insights from the LLMs' choice behavior and chain-of-thought, (3) validate our litmus test framework by testing self-consistency, robustness, and generalizability. Overall, this work provides a foundation for evaluating LLM agents as they are further integrated into economic decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。