arXiv:2503.16743cs.AIcs.IT2025-03

用算法信息论测试超智能,发现大模型预测力与压缩能力直接相关。

Can Complexity and Uncomputability Explain Intelligence? SuperARC: A Test for Artificial Super Intelligence Based on Recursive Compression

  • 基于算法信息论设计无人类依赖的复杂性评测体系
  • 顶尖大模型在压缩任务上表现优于多数模型但版本间波动明显
  • 符号与神经融合方法在抽象压缩中超越纯统计模型

我们提出一种递增复杂度、开放性且不依赖人类的评估指标,用于检验人工智能在通用智能(AGI)与超智能(ASI)宣称下的表现。该测试基于随机性与最优推断等数学基础,不依赖人类提问或模式匹配。我们主张,基于算法信息论(AIT)建立的普适原则,可为模型抽象与预测提供强有力的度量框架。在前沿模型测试中,领先的大语言模型在多项任务中表现优异,但其最新版本常出现退化,未逼近由AIT定义的通用智能(UAI)理论峰值。相反,基于相同原则的混合神经-符号方法在压缩驱动的模型抽象与序列预测中,优于专门的预测模型。最终证明:对任意形式理论的预测能力,与算法空间的压缩程度成正比,而非统计空间。因此,未来模型进展必须结合符号方法,而当前大模型开发者往往忽视或未意识到这一点。

原文摘要 · Abstract (English)

We introduce an increasing-complexity, open-ended, and human-agnostic metric to evaluate foundational and frontier AI models in the context of Artificial General Intelligence (AGI) and Artificial Super Intelligence (ASI) claims. Unlike other tests that rely on human-centric questions and expected answers, or on pattern-matching methods, the test here introduced is grounded on fundamental mathematical areas of randomness and optimal inference. We argue that human-agnostic metrics based on the universal principles established by Algorithmic Information Theory (AIT) formally framing the concepts of model abstraction and prediction offer a powerful metrological framework. When applied to frontiers models, the leading LLMs outperform most others in multiple tasks, but they do not always do so with their latest model versions, which often regress and appear far from any global maximum or target estimated using the principles of AIT defining a Universal Intelligence (UAI) point and trend in the benchmarking. Conversely, a hybrid neuro-symbolic approach to UAI based on the same principles is shown to outperform frontier specialised prediction models in a simplified but relevant example related to compression-based model abstraction and sequence prediction. Finally, we prove and conclude that predictive power through arbitrary formal theories is directly proportional to compression over the algorithmic space, not the statistical space, and so further AI models' progress can only be achieved in combination with symbolic approaches that LLMs developers are adopting often without acknowledgement or realisation.

超智能算法信息论大模型评测符号推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。