用量子物理方法在稀疏图上嵌入图像特征,显著提升分类精度且参数量极低。
Kohn-Sham Spectral Embedding on Sparse Graphs at the Nishimori Temperature for Image Classification

- 基于随机键伊辛模型的奈希莫里温度,将特征映射到稀疏图进行谱嵌入
- 在ImageNet-1000上达88.93%准确率,参数仅2124万,远少于Swin-L和ViT-H/14
- 适合追求高效高精度模型的开发者,尤其关注小参数量下的性能突破
我们提出基底-沙姆谱嵌入(KSSE),一种能量模型,用稀疏图谱嵌入替代卷积网络顶层分类器。该嵌入位于关联随机键伊辛模型的奈希莫里温度下,即类结构与无序勉强可区分的谱检测阈值。将预训练特征映射至准循环低密度奇偶校验图,构建正则化拉普拉斯矩阵(贝斯-赫森矩阵)作为有效基底-沙姆哈密顿量,每通道独立求解,通过傅里叶变换在循环块上以$O(N \log N + k_{\text{mode}}^{2} N)$时间完成,辅以低模态瑞利-里茨修正($k_{\text{mode}}=5$)。物理上等价于一维环晶格上的k.p有效质量降低:循环支撑为完美晶体,数据权重为缓慢变化的杂质势,奈希莫里交叉对应能带边缘的费米能级。星域手术优化图结构:不消除所有受挫环(否则破坏码字),而是通过边移位在码字周围建立认证凸性,残余受挫有界,实现多尺度分形认证(分形维数$D_{2}<1$ vs 粗糙景观$D_{2}>3$)。理论包含广义伊哈-巴斯恒等式、尖锐谱阈值、非回溯增长三态、受挫作为规范不变$Z_{2}$通量、陷阱集谱检验、精确通道可分离性及杯积障碍,以及环级数、凸性、手术与准稳态边界。在冻结的EfficientNet-B4特征($D=1792$)下,转导协议中,KSSE在ImageNet-1000上达到88.93%的Top-1准确率,约2124万参数,超越Swin-L(197M,86.4–87.3%),媲美ViT-H/14(632M,88.0–89.5%)但参数减少10倍与30倍。
原文摘要 · Abstract (English)
We propose Kohn-Sham Spectral Embedding (KSSE), an energy-based model replacing the top-layer classifier of convolutional networks with a sparse-graph spectral embedding at the Nishimori temperature of an associated Random-Bond Ising Model the spectral detectability threshold where class structure becomes marginally distinguishable from disorder. Mapping pre-trained features onto quasi-cyclic low-density parity-check graphs, we construct a regularized Laplacian (Bethe-Hessian) as an effective Kohn-Sham Hamiltonian, yielding D independent spectral problems-one per feature channel-solvable in $O(N log N + k_{mode}^{2} N)$ time by FFT on circulant blocks (Pontryagin self-duality), with low-mode Rayleigh-Ritz refinement ($k_{mode}=5$). Physically, this is a k.p effective-mass reduction on a one-dimensional ring crystal: the circulant support is the perfect crystal, the data weights a slowly varying impurity potential, and the Nishimori crossing a Fermi level at the band edge. Star-domain surgery optimizes the graph: instead of eliminating all frustrated cycles impossible without destroying the codewords-edge shifts create certified convexity around codewords with bounded residual frustration, with multi-scale fractal certification (basins $D_{2}<1$ vs rough landscapes $D_{2}>3$). The theory includes a generalized Ihara-Bass identity with a sharp spectral threshold, a non-backtracking growth trichotomy with frustration as a gauge-invariant $Z_{2}$ flux, a trapping-set spectral test, exact channel separability with a cup-product obstruction, plus loop-series, convexity, surgery, and quasi-stationarity bounds. On ImageNet-1000 with frozen EfficientNet-B4 features (D=1792) under a transductive protocol, KSSE achieves 88.93% Top-1 accuracy with ~21.24M parameters-beating Swin-L (197M, 86.4-87.3%) and matching the lower end of ViT-H/14 (632M, 88.0-89.5%) with 10x and 30x fewer parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。