arXiv:2505.13511cs.AIcs.CL2025-05被引 2

测试AI在真实自由职业编程任务中的赚钱能力,发现顶级模型年入百万美元。

Can AI Freelancers Compete? Benchmarking Earnings, Reliability, and Task Success at Scale

  • 用合成任务构建可自动评测的自由职业基准,支持大规模重复测试。
  • Claude 3.5 Haiku表现最佳,累计赚取约152万美元,远超其他模型。
  • 适合关注AI自动化开发可行性及评估方法的研究者与从业者。

本研究探讨大型语言模型(LLMs)作为自主代理完成真实世界任务的潜力,包括自由职业软件开发。本文构建了一个新基准,基于Kaggle自由职业数据集生成合成任务,涵盖编程与数据分析任务,所有任务价格标准化为美元(中位数固定项目价格约250美元,平均306美元)。每个任务配备结构化输入输出测试用例和预估价格标签,支持自动化正确性验证与货币化绩效评估。该框架受OpenAI SWE-Lancer基准启发(1,400个真实Upwork任务,总价值100万美元),但通过程序化可测任务与预测价格实现更高可扩展性与可复现性。在该基准上评估了Claude 3.5 Haiku、GPT-4o-mini、Qwen 2.5和Mistral四个现代LLM,报告其准确率(任务成功率与测试用例通过率)及累计“自由职业收入”(已解决问题的总价)。结果显示,Claude 3.5 Haiku表现最优,收入约152万美元,紧随其后的是GPT-4o-mini(149万美元)、Qwen 2.5(133万美元)和Mistral(70万美元)。分析显示,最强模型解决任务最多且极少完全失败。论文讨论了该结果对AI作为自由职业开发者可行性的启示,以及自动化评估方法的优势与局限,并指出结构化任务表现与真实复杂性之间的差距。

原文摘要 · Abstract (English)

This study explores Large Language Models (LLMs) as autonomous agents for real-world tasks, including freelance software development. This work presents a new benchmark that evaluates LLMs on freelance programming and data analysis tasks derived from economic data. We construct the benchmark using synthetic tasks created from a Kaggle Freelancer dataset of job postings, with all job prices standardized to USD (median fixed-project price around $250, and an average of $306). Each task is accompanied by structured input-output test cases and an estimated price tag, enabling automated correctness checking and a monetary performance valuation. This approach is inspired by OpenAI's recent SWE-Lancer benchmark (1,400 real Upwork tasks worth $1M total). Still, our framework simplifies evaluation using programmatically testable tasks and predicted price values, making it highly scalable and repeatable. On this benchmark, we evaluate four modern LLMs - Claude 3.5 Haiku, GPT-4o-mini, Qwen 2.5, and Mistral. We report each model's accuracy (task success rate and test-case pass rate) and the total "freelance earnings" it achieves (sum of prices of solved tasks). Our results show that Claude 3.5 Haiku performs best, earning approximately $1.52 million USD, followed closely by GPT-4o-mini at $1.49 million, then Qwen 2.5 ($1.33M) and Mistral ($0.70M). We analyze the distribution of errors per task and observe that the strongest models solve the most tasks and rarely fail completely on any project. We discuss the implications of these results for the feasibility of AI as a freelance developer, the advantages and limitations of our automated benchmark approach, and the gap between performance on structured tasks versus the true complexity of real-world freelance jobs.

AI自由职业大模型评估自动化测试任务基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。