arXiv:2607.12634cs.AI2026-07

智能的本质是信息压缩与重组,用原子单元提升系统可扩展性。

Atomic Units of X: The Compression Layer of Intelligence

  • 提出压缩演算框架,量化信息如何通过原子单元重组
  • 在大型软件项目中实现显著的概念整合,小项目未达标
  • 揭示大模型本质是动态组合结构而非知识存储库

本文提出一个理论与实证框架,将智能视为原子压缩与组合复用的过程。可扩展的认知、生物、计算及组织系统通过将信息组织为可复用的原子单元,实现复杂度降低,并可重组为更高阶结构。核心贡献为压缩演算(Compression Calculus),用于比较表面证据与原子表征,描述抽象在多层间的累积效应。该框架在大规模多源软件语料库上基于开放世界概念模型进行评估,显示两个大型项目中证据到概念的整合程度显著提升,而第三个小项目未达目标阈值。分析进一步识别出关键边界条件:整合效果依赖语料规模、证据单元定义与概念身份;观察到的减少主要源于语料内重复,而非跨源重复。论文还构建了对当代生成系统意义鸿沟的表征解释,将潜在的、上下文相关的概念近似视为缺乏持久身份与组合约束的软原子,从而提出一种架构视角:大语言模型作为动态融合引擎,负责导航并组合稳定的概念结构,而非单一知识存储。

原文摘要 · Abstract (English)

This paper proposes a theoretical and empirical framework for understanding intelligence as a process of atomic compression and compositional reuse. It argues that scalable cognitive, biological, computational, and organisational systems reduce complexity by organising information into reusable units that can be recombined into higher-order structures. The central contribution is the Compression Calculus, a formal framework for comparing surface evidence with atomic representations and for describing how abstraction can compound across layers. The framework is evaluated on large, multi-source software corpora under an open-world concept model, showing substantial evidence-to-concept consolidation in two large projects, while a smaller third project remains below the target threshold. The analysis further identifies important boundary conditions: consolidation depends on corpus scale, evidence-unit definition, and concept identity, and the observed reduction is driven primarily by within-corpus recurrence rather than cross-source recurrence. The paper also develops a representational account of the meaning gap in contemporary generative systems, describing latent, context-dependent conceptual approximations as soft atoms that lack the persistent identity and compositional constraints of stable atomic units. This motivates an architectural view in which large language models function as dynamic fusion engines that navigate and compose persistent conceptual structures rather than serving as the sole repository of those structures.

认知科学压缩演算大模型架构概念抽象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。