arXiv:2502.12115cs.LGcs.SE2025-02ICML被引 113

测试大模型在真实软件外包任务中的赚钱能力,发现仍难胜任多数工作。

SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

  • 构建1400+真实外包任务库,涵盖从50美元到3.2万美元的开发与管理任务。
  • 前沿模型仅能完成少数任务,独立任务需三重工程师验证,管理决策比对原项目经理选择。
  • 开源完整评估环境,助力研究AI对软件经济的影响。

我们提出SWE-Lancer,一个包含超过1400个来自Upwork的真实软件工程自由职业任务的基准,总价值达100万美元。任务包括独立开发(从50美元的漏洞修复到3.2万美元的功能实现)和管理决策类任务,后者要求模型在多个技术方案间做出选择。独立任务通过三重验证的端到端测试评分,管理决策则以原始雇佣经理的选择为金标准进行评估。实验表明,当前前沿大模型仍无法解决多数任务。为促进后续研究,我们开源统一Docker镜像及公开评估集SWE-Lancer Diamond(https://github.com/openai/SWELancer-Benchmark)。通过将模型表现映射至实际收益,我们希望推动对人工智能发展经济影响的深入研究。

原文摘要 · Abstract (English)

We introduce SWE-Lancer, a benchmark of over 1,400 freelance software engineering tasks from Upwork, valued at \$1 million USD total in real-world payouts. SWE-Lancer encompasses both independent engineering tasks--ranging from \$50 bug fixes to \$32,000 feature implementations--and managerial tasks, where models choose between technical implementation proposals. Independent tasks are graded with end-to-end tests triple-verified by experienced software engineers, while managerial decisions are assessed against the choices of the original hired engineering managers. We evaluate model performance and find that frontier models are still unable to solve the majority of tasks. To facilitate future research, we open-source a unified Docker image and a public evaluation split, SWE-Lancer Diamond (https://github.com/openai/SWELancer-Benchmark). By mapping model performance to monetary value, we hope SWE-Lancer enables greater research into the economic impact of AI model development.

大模型软件工程真实任务经济评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。