开源指数量化AI在各职业的采用率与执行能力,发现金融、计算机、艺术行业采用最高。
The Open Source Economic Index of AI Adoption and Capability

- 基于公开对话数据和O*NET任务构建开源经济指数
- 测试显示AI能完成高层级任务但细节常出错
- 适合关注AI落地现状与局限的研究者或政策制定者
我们致力于衡量AI在不同职业中对离散劳动任务的采纳程度及其执行能力。为衡量采纳率,开发了一个开源经济指数,利用公开可用的用户-大模型聊天数据与O*NET任务,复现前沿AI实验室的研究成果,发现金融、计算机科学和艺术行业的采纳率最高。为衡量能力,构建了一套基于O*NET职业、任务和模型-上下文-协议(MCP)服务器的基准场景系统。在9个指数中高频出现的职业上,使用Kimi-k2.5与OpenAI代理SDK测试,结果显示AI能正确执行高层级工作流,但在具体工具调用等细粒度操作上常出现错误。
原文摘要 · Abstract (English)
We work towards measuring both AI adoption and the capability of AI to perform discrete labor tasks across various occupations. To measure adoption, we develop an open-source economic index that uses publicly available user-LLM chat data and O*NET tasks to replicate studies produced by frontier AI labs, finding that occupations in the finance, computer science, and arts sectors are those with the highest adoption rates. To measure capabilities, we build a system that generates benchmark scenarios grounded in O*NET occupations, tasks, and model-context-protocol (MCP) servers. We test Kimi-k2.5 with an OpenAI agents SDK harness on scenarios across 9 occupations that appear frequently in our index, finding that AI correctly executes high-level workflows but often errs in the granular details (such as specific tool calls used).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。