研究随机矩阵的多样性,为量子方程的上下文学习提供理论保障
A Theory of Diversity for Random Matrices with Applications to In-Context Learning of Schrödinger Equations
- 分析多组随机矩阵的中心化子是否平凡,推导概率下界
- 发现当样本数足够大时,中心化子几乎必然平凡
- 适用于基于Transformer的量子系统学习,理论意义强
我们研究如下问题:给定一组从共同分布 𝕡 独立抽取的 d×d 随机矩阵 {A^(1),…,A^(N)},其生成的代数中心化子为平凡的概率是多少?针对由随机势能离散化线性薛定谔算子产生的几类随机矩阵,我们给出了该概率关于样本数 N 与维度 d 的下界。结合近期机器学习理论成果,我们的结果为基于Transformer的神经网络在薛定谔方程的上下文学习中的泛化能力提供了理论保证。
原文摘要 · Abstract (English)
We address the following question: given a collection $\{\mathbf{A}^{(1)}, \dots, \mathbf{A}^{(N)}\}$ of independent $d \times d$ random matrices drawn from a common distribution $\mathbb{P}$, what is the probability that the centralizer of $\{\mathbf{A}^{(1)}, \dots, \mathbf{A}^{(N)}\}$ is trivial? We provide lower bounds on this probability in terms of the sample size $N$ and the dimension $d$ for several families of random matrices which arise from the discretization of linear Schrödinger operators with random potentials. When combined with recent work on machine learning theory, our results provide guarantees on the generalization ability of transformer-based neural networks for in-context learning of Schrödinger equations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。