arXiv:2607.18433cs.LGcs.AI2026-07

提出可学习的新奇性,统一智能的多个表现形式。

Intelligence from Learnable Novelty

论文配图:Intelligence from Learnable Novelty
图 1 · 摘自论文原文
  • 定义可学习新奇性,分离可转化为知识的意外信息
  • 无需监督即可准确排序元胞自动机复杂度,最高达规则110
  • 在无监督图像分类和强化学习探索中表现优异

智能在不同领域以不同形式出现:统计与机器学习中的数据压缩、动力系统中的通用计算、以及智能体中的适应性行为。各领域目标各异,但两大主流方法常陷入镜像困境:新颖性搜索追求意外,却困于噪声电视屏幕;自由能原理避免意外,却最满足于黑暗房间。两者失败根源在于:将学习者可转化的知识性意外与不可转化的意外混为一谈。本文提出‘可学习新奇性’概念,其可解析估计,基于廉价且可微分的储层计算机实现。作为无监督度量,该估计器成功复现数十年的复杂性分类,将图灵完备规则110排在初等元胞自动机之首。作为优化目标,其梯度引导神经元胞自动机从简单动态演化至孤子态——规则110进行计算的传播结构;同时,在不使用任何标签的情况下,组织图像编码器对MNIST十类数字的表示。作为内在奖励引入强化学习代理,弥补任务奖励的探索不足,在十个环境中有九个超越基线,且无崩溃。复杂性生成、抽象与探索,原本由不同目标独立驱动的三类智能行为,如今均源自对单一可微量的上升优化,使智能的多重投影获得统一量化基础。

原文摘要 · Abstract (English)

Intelligence appears under different names in different fields: as data compression in statistics and machine learning, as universal computation in dynamical systems, and as adaptive behavior in agents. Each field carries its own objective, and the two most influential drives often fail in mirror image: novelty search, which seeks surprise, is transfixed by a noisy television screen, while the free-energy principle, which avoids surprise, is most content in a dark room. Both failures have a single cause: each objective treats as one quantity the surprise a learner can convert into knowledge and the surprise it never can. Here we show that the learnable part of that information, which we call learnable novelty, yields the seemingly disparate projections of intelligence, and we give a closed-form estimator of it built on a cheap and differentiable reservoir computer. Used as a measure, with no supervision of any kind, the estimator recovers decades of complexity classification, ranking the Turing-complete rule~110 highest among the elementary cellular automata. Used as an objective, its gradient carries a neural cellular automaton from simple dynamics into a regime of solitons, the traveling, colliding structures by which rule~110 computes, as well as organizes the representation of an image encoder around the ten digit classes of MNIST, fully unsupervised: no label ever enters training. Handed to a reinforcement-learning agent as an intrinsic reward, it supplies the exploration that task rewards lack, improving on the task baseline in nine of ten environments and collapsing in none. Complexity generation, abstraction, and exploration, ordinarily pursued with unrelated objectives in separate fields, thus emerge from ascent on one differentiable quantity, and the projections of intelligence gain a common quantitative footing.

智能统一无监督学习复杂性强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。