arXiv:2502.18858cs.AIcs.CL2025-02被引 2

用试错失败次数评估智能,发现当前AI在复杂任务上仍远未达标。

Evaluating Intelligence via Trial and Error

  • 以试错失败次数衡量智能,失败越少代表智能越高。
  • 现有AI在简单任务达自主水平,复杂任务如语言、视觉仍差距巨大。
  • 人类任务具临界性特征,当前AI依赖表面模仿而非深层理解。

智能是物种在有限试错中找到解决方案的关键能力。基于此,我们提出生存游戏(Survival Game)框架,通过试错过程中的失败次数评估智能:失败越少,智能越高。当失败次数的期望和方差均有限时,表明系统能稳定应对新挑战,定义为自主水平(Autonomous Level)。我们全面评估现有AI系统,结果表明:尽管在简单任务中已达到自主水平,但在视觉、搜索、推荐和语言等复杂任务中仍相距甚远。若仅靠当前技术扩展,实现通用任务的自主水平需约10^26个参数。这相当于所需H100 GPU总价值为苹果公司市值的10^7倍,即使按摩尔定律推进,也需70年才能支撑该参数规模。理论分析揭示人类任务具有临界性,达成自主水平需深刻理解任务机制。当前AI缺乏深层理解,仅依赖表面模仿,难以真正自主。我们认为生存游戏不仅可引导AI未来发展,亦能深化对人类智能的理解。

原文摘要 · Abstract (English)

Intelligence is a crucial trait for species to find solutions within a limited number of trial-and-error attempts. Building on this idea, we introduce Survival Game as a framework to evaluate intelligence based on the number of failed attempts in a trial-and-error process. Fewer failures indicate higher intelligence. When the expectation and variance of failure counts are both finite, it signals the ability to consistently find solutions to new challenges, which we define as the Autonomous Level of intelligence. Using Survival Game, we comprehensively evaluate existing AI systems. Our results show that while AI systems achieve the Autonomous Level in simple tasks, they are still far from it in more complex tasks, such as vision, search, recommendation, and language. While scaling current AI technologies might help, this would come at an astronomical cost. Projections suggest that achieving the Autonomous Level for general tasks would require $10^{26}$ parameters. To put this into perspective, loading such a massive model requires so many H100 GPUs that their total value is $10^{7}$ times that of Apple Inc.'s market value. Even with Moore's Law, supporting such a parameter scale would take $70$ years. This staggering cost highlights the complexity of human tasks and the inadequacies of current AI technologies. To further investigate this phenomenon, we conduct a theoretical analysis of Survival Game and its experimental results. Our findings suggest that human tasks possess a criticality property. As a result, Autonomous Level requires a deep understanding of the task's underlying mechanisms. Current AI systems, however, do not fully grasp these mechanisms and instead rely on superficial mimicry, making it difficult for them to reach an autonomous level. We believe Survival Game can not only guide the future development of AI but also offer profound insights into human intelligence.

智能评估试错学习人工智能局限自主性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。