提出新测试框架,判断AI是否具备通用智能
Turing Test 2.0: The General Intelligence Threshold
- 定义通用智能阈值,明确区分是否达AGI
- 设计可清晰判定通过/失败的测试机制
- 适用于当前大模型,提供实操评估方法
随着人工智能和大型语言模型(如ChatGPT)的发展,实现人工通用智能(A.G.I.)的新竞赛已开启。尽管各方对A.G.I.何时达成存在猜测,但目前尚无明确标准来检测模型是否真正达到或超越通用智能,即便使用传统的图灵测试及其现代变体也难以有效衡量。本文指出传统方法在检测A.G.I.上的不足,并提出两项新贡献:第一,给出通用智能(G.I.)的明确定义,并设立可量化的通用智能阈值(G.I.T.),用于区分是否达到A.G.I.;第二,构建一套新框架,指导如何设计能以简单、全面且明确的“通过/失败”方式检测系统是否具备通用智能的测试。该框架被称为图灵测试2.0。文中还展示了将此框架应用于现代AI模型的实际案例。
原文摘要 · Abstract (English)
With the rise of artificial intelligence (A.I.) and large language models like ChatGPT, a new race for achieving artificial general intelligence (A.G.I) has started. While many speculate how and when A.I. will achieve A.G.I., there is no clear agreement on how A.G.I. can be detected in A.I. models, even when popular tools like the Turing test (and its modern variations) are used to measure their intelligence. In this work, we discuss why traditional methods like the Turing test do not suffice for measuring or detecting A.G.I. and provide a new, practical method that can be used to decide if a system (computer or any other) has reached or surpassed A.G.I. To achieve this, we make two new contributions. First, we present a clear definition for general intelligence (G.I.) and set a G.I. Threshold (G.I.T.) that can be used to distinguish between systems that achieve A.G.I. and systems that do not. Second, we present a new framework on how to construct tests that can detect if a system has achieved G.I. in a simple, comprehensive, and clear-cut fail/pass way. We call this novel framework the Turing test 2.0. We then demonstrate real-life examples of applying tests that follow our Turing test 2.0 framework on modern A.I. models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。