大脑与深度网络的相似表征源于对信息压缩的共同追求。
The Information-Theoretic Imperative: Compression and the Epistemic Foundations of Intelligence
- 用压缩效率原理解释神经表征如何选择稳定不变量
- 预测了不变表征在分布外任务中更鲁棒,且存在主导阈值
- 适合研究智能本质、神经编码或模型泛化性的学者
为何大脑与深度网络会发展出相似的表征?尽管基质和优化路径完全不同,任务优化的神经网络仍能精确预测灵长类腹侧流的响应。这种收敛不能仅由自然图像统计或任务结构解释。压缩效率原理(CEP)提出:利用不稳定的关联会产生不断增长的“例外税”(约线性增加的编码长度),而编码转移稳定的不变量则可分摊此成本。当环境提供丰富的干预性变化且具有近似模块化的因果结构时,这些不变量恰好对应因果机制。该框架统一解释了三个生物特征——神经信号的高代谢成本、早期感官通路的高编码效率、腹侧流的层级容错性——并将其与深度学习中的现象对应:规模极限、分布外时的捷径失效,以及数据增强对不变性的强制作用。关键预测包括:不变表征占优的交叉阈值,以及压缩效率与分布外鲁棒性的系统耦合——可在不同基质上验证。生物稀疏信号与深度网络过参数化的差异,源于对同一权衡拓扑的不同资源约束。这种收敛并非偶然,而是由在变化下进行预测性压缩所驱动的跨基质极小值盆地的证据。
原文摘要 · Abstract (English)
Why do brains and deep networks converge on similar representations? Task-optimized artificial neural networks quantitatively predict primate ventral stream responses despite radically different substrates and optimization dynamics. This convergence demands explanation beyond shared natural image statistics or task structure alone. The Compression Efficiency Principle (CEP) specifies the selection mechanism: representations exploiting unstable correlations pay a growing "exception tax" (approximately linear excess codelength under shortcut-flipping shifts), while representations encoding shift-stable invariants amortize this cost. When environments provide intervention-rich shifts and exhibit approximately modular causal structure, these invariants align with causal mechanisms. The framework offers a unified lens on three biological signatures -- steep metabolic constraints on neural signaling, high coding efficiency in early sensory pathways, and hierarchical tolerance in the ventral stream -- and connects them to parallel phenomena in deep learning: scaling frontiers, shortcut failures under distribution shift, and the role of augmentation in enforcing invariances. Distinctive predictions follow: a crossover threshold beyond which invariant representations dominate, and systematic coupling between compression efficiency and out-of-distribution robustness -- testable across substrates. Predicted divergences (sparse biological signaling versus dense overparameterization) arise from different resource constraints on a shared trade-off topology. The convergence is not a coincidence. It is evidence for a substrate-independent basin shaped by predictive compression under shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。