arXiv:2504.15434cs.AIcs.CV2025-04

测试AI玩猜词游戏,发现其表现远低于人类。

AGI Is Coming... Right After AI Learns to Play Wordle

  • 用计算机界面控制的AI代理玩《纽约时报》猜词游戏
  • 连续多日测试,正确率仅5.36%
  • 揭示当前大模型在简单任务中仍存严重缺陷

本文研究了多模态智能体,特别是OpenAI的Computer-User Agent(CUA),该代理通过标准计算机界面执行任务,类似于人类操作。我们通过《纽约时报》的Wordle游戏评估其表现,以揭示模型行为并识别不足之处。结果显示,模型在不同情境下对颜色的识别能力存在显著差异。在数天内多次运行测试中,成功率仅为5.36%。尽管人工智能代理引发广泛期待,并可能推动人工通用智能(AGI)的到来,但我们的结果表明,即便是简单的任务,对当前前沿AI模型而言仍是巨大挑战。文章最后讨论了潜在原因、对未来发展的启示及改进方向。

原文摘要 · Abstract (English)

This paper investigates multimodal agents, in particular, OpenAI's Computer-User Agent (CUA), trained to control and complete tasks through a standard computer interface, similar to humans. We evaluated the agent's performance on the New York Times Wordle game to elicit model behaviors and identify shortcomings. Our findings revealed a significant discrepancy in the model's ability to recognize colors correctly depending on the context. The model had a $5.36\%$ success rate over several hundred runs across a week of Wordle. Despite the immense enthusiasm surrounding AI agents and their potential to usher in Artificial General Intelligence (AGI), our findings reinforce the fact that even simple tasks present substantial challenges for today's frontier AI models. We conclude with a discussion of the potential underlying causes, implications for future development, and research directions to improve these AI systems.

AI代理语言理解认知缺陷游戏测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。