神经网络能自发学习量子与后量子生成模型,像贝叶斯更新一样推断未来。
Neural networks leverage nominally quantum and post-quantum representations
- 用标准预训练方式让网络发现隐含的量子/后量子生成机制。
- 激活空间中的几何关系与网络架构无关,体现历史对未来的概率影响。
- 适合研究模型内在表征、认知建模或量子启发计算的读者。
我们发现,深度神经网络(包括Transformer和RNN)在常规的下一个词预测任务上预训练后,会内在地发现并表示其训练数据的‘量子’与‘后量子’低维生成模型——仿佛在推理时随着上下文观察不断对世界模型的潜在状态进行迭代贝叶斯更新。值得注意的是,神经网络能轻松实现这一表征,而任何有限经典电路都无法完成该任务。不同输入序列引起的神经激活间几何关系,几乎与网络架构无关。该几何空间中每一点对应一个由历史诱导的未来概率密度,点之间的相对位移反映了不同过往对未来的机制与强度差异。
原文摘要 · Abstract (English)
We show that deep neural networks, including transformers and RNNs, pretrained as usual on next-token prediction, intrinsically discover and represent beliefs over 'quantum' and 'post-quantum' low-dimensional generative models of their training data -- as if performing iterative Bayesian updates over the latent state of this world model during inference as they observe more context. Notably, neural nets easily find these representation whereas there is no finite classical circuit that would do the job. The corresponding geometric relationships among neural activations induced by different input sequences are found to be largely independent of neural-network architecture. Each point in this geometry corresponds to a history-induced probability density over all possible futures, and the relative displacement of these points reflects the difference in mechanism and magnitude for how these distinct pasts affect the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。