arXiv:2601.03220cs.LGstat.ML2026-01被引 37

提出新信息度量方法,揭示计算受限智能体如何从数据中构造有用信息。

From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence

  • 引入'表征复杂度'概念,衡量计算有限者能从数据中学到什么
  • 证明通过计算可创造新信息,且数据顺序影响可学习内容
  • 可用于指导数据筛选与生成,提升模型泛化能力

现有信息论(香农与科尔莫戈罗夫)在面对计算能力受限的观察者时失效,无法解释为何某些数据变换能产生新知识。本文指出三大悖论:(1)确定性变换不能增加信息;(2)信息不依赖数据顺序;(3)似然建模只是分布匹配。为破解矛盾,提出“表征复杂度”(epiplexity),量化计算受限智能体可提取的结构化信息,排除时间受限熵(如伪随机数生成器中的不可预测成分)。实验表明,通过计算可生成超越原始过程的新信息,数据顺序显著影响可学内容,且似然建模能生成比原数据更复杂的程序。提出实用估计方法,能区分不同数据源、跟踪下游性能,并识别提升跨分布泛化能力的数据干预策略。该框架为数据选择与生成提供理论依据,突破传统模型选择范式。

原文摘要 · Abstract (English)

Can we learn more from data than existed in the generating process itself? Can new and useful information be constructed from merely applying deterministic transformations to existing data? Can the learnable content in data be evaluated without considering a downstream task? On these questions, Shannon information and Kolmogorov complexity come up nearly empty-handed, in part because they assume observers with unlimited computational capacity and do not target the useful information content. In this work, we identify and exemplify three seeming paradoxes in information theory: (1) information cannot be increased by deterministic transformations; (2) information is independent of the order of data; (3) likelihood modeling is merely distribution matching. To shed light on the tension between these results and modern practice, and to quantify the value of data, we introduce epiplexity, a formalization of information capturing what computationally bounded observers can learn from data. Epiplexity captures the structural content in data while excluding time-bounded entropy, the random unpredictable content exemplified by pseudorandom number generators and chaotic dynamical systems. With these concepts, we demonstrate how information can be created with computation, how it depends on the ordering of the data, and how likelihood modeling can produce more complex programs than present in the data generating process itself. We also present practical procedures to estimate epiplexity which we show capture differences across data sources, track with downstream performance, and highlight dataset interventions that improve out-of-distribution generalization. In contrast to principles of model selection, epiplexity provides a theoretical foundation for data selection, guiding how to select, generate, or transform data for learning systems.

信息论数据选择计算受限表征复杂度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。