arXiv:2605.23920cs.CYcs.AI2026-05

AI已能低成本完成多数真实努力任务,动摇实验经济学基础。

Artificial Effort

论文配图:Artificial Effort
图 1 · 摘自论文原文
  • 用23个大模型测试8类真实努力任务,多数可被准确自动化。
  • 模型越新表现越好,中端模型快速逼近顶尖水平。
  • 给钱激励对AI无效,适合关注人机交互与实验可信度的研究者。

真实努力任务(即参与者需付出认知成本且结果取决于实际表现的任务)在实验经济学中广泛应用。其有效性依赖于人类亲自完成的前提。我们考察在人工智能与大型语言模型(LLMs)时代,这一前提是否仍成立。基于8个经典真实努力任务和来自三大厂商的23个LLM,研究发现:多数任务可被准确且以可忽略成本解决,仅有少数难以自动化。模型性能随代际提升,中端模型正迅速缩小与前沿模型的差距,使更多通用模型具备自动化能力。此外,口头提供货币激励对LLM表现无影响。结果表明,在无监督场景下,若参与者可廉价外包任务给LLM,观察到的表现可能不再反映真实人类努力,从而确立了真实努力任务使用的边界条件。

原文摘要 · Abstract (English)

Real-effort tasks, in which participants perform cognitively costly activities whose outcomes depend on actual performance, are widely used in experimental economics. Their validity, however, rests on the assumption that a human performs them. We study whether this assumption still holds in the era of Artificial Intelligence (AI) and Large Language Models (LLMs). Using 8 canonical real-effort tasks and 23 LLMs from three major providers, we show that most tasks can now be solved accurately and at a negligible cost, while only a few resist automation. Performance improves with each model generation, and midtier models are rapidly closing the gap with frontier ones, broadening the set of widely accessible models that can automate these tasks. Additionally, we show that verbally offering monetary incentives has no effect on LLM performance. Our findings establish a boundary condition for the use of real-effort tasks in unsupervised settings: when participants can cheaply outsource task completion to an LLM, observed performance may no longer reflect genuine human effort.

AI评估实验经济学语言模型自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。