评估大模型在真实工作中的经济价值,发现其能节省大量时间但使用率远低于潜力。
Economic Evaluations of Language Models

- 基于真实用户查询和合成数据构建评估体系,覆盖47%美国职业
- 当前模型可为50%以上任务节省时间,79%高潜力任务实际使用率低
- 揭示隐私与专有系统是限制应用的核心瓶颈,适合政策与企业决策者
语言模型执行具有经济价值的任务,但尚未系统评估其在各类经济活动中的表现。本文提出EconEvals——一个开源评估套件,用于衡量模型在真实美国劳动力市场任务、工作活动及职业中的能力。评估基于真实用户查询,并补充合成数据,覆盖范围超越OpenAI的GDPval基准(原覆盖5%职业),成本降低500倍。同时引入基于模拟的暴露度量,估算当前模型能力可节省全美所有职业任务的总时间,详细列出每项估计。结果显示,当前模型可在至少一半任务中为47%的职业节省大量时间;然而,在79%预计显著节省时间的任务中,实际使用量(如Claude)仍很低,表明使用远落后于潜力。除模型聊天机器人固有限制外,隐私顾虑和专有系统是主要障碍。整体框架具备可扩展性,能随模型能力迭代持续更新,为评估语言模型对劳动力市场的实际影响提供可落地的基础设施。
原文摘要 · Abstract (English)
Language models perform economically valuable work, yet they are not currently assessed for how well they perform every economically valuable task. We introduce EconEvals as an open-source evaluation suite to measure capabilities relevant to tasks, work activities, and occupations in the US labor economy. We ground the evaluation suite in real user queries to language models where possible, and supplement these with synthetic data. Our evaluations improve coverage over OpenAI's GDPval benchmark, which is the existing state-of-the-art that covers 5% of US occupations, at 500x lower cost. Alongside benchmarks, we also introduce a simulation-based exposure measure to estimate how much time current language model capabilities could save across all tasks belonging to all US occupations, with detailed accounting for each estimate. Our estimates indicate that current models could save workers substantial time on at least half of their tasks in 47% of occupations. However, for 79% of tasks where we predict substantial time savings, observed Claude usage is low, suggesting that existing usage lags potential. Beyond inherent constraints of language model chatbots, our data identifies privacy and proprietary systems as the principal bottlenecks limiting further time savings from AI. Overall, we introduce adaptable infrastructure that grounds inferences about language models' labor-market impact in their current capabilities, which can be continually updated as capabilities improve.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。