arXiv:2508.18717cs.LGcs.CV2025-08中稿 · the Moscow Univers…被引 2

用物理模型压缩图像特征,实现高效高精度分类。

Natural Image Classification via Quasi-Cyclic Graph Ensembles and Random-Bond Ising Models at the Nishimori Temperature

  • 将图像特征转为物理模型中的自旋,构建分层图结构进行分类。
  • 在ImageNet-10和ImageNet-100上分别达到98.7%和84.92%准确率。
  • 适合追求低功耗、高效率模型的部署场景。

现代多类图像分类依赖高维CNN特征,带来巨大内存与计算开销,并遮蔽数据流形几何结构。现有基于图的谱分类器适用于合成或二分类任务,但在自然图像多类别场景中表现下降,因特征流形具有非平凡拓扑。本文提出一种受物理启发的流程:将冻结的MobileNetV2特征视为稀疏多边类型准循环LDPC图上的伊辛自旋,定义随机键伊辛模型(RBIM)。模型在尼西莫里温度下运行——此时贝斯-海森矩阵最小特征值为零。谱-拓扑对应关系将托纳图中的捕获集与代数拓扑缺陷通过伊哈-巴思ζ函数极点关联,系统抑制有害子结构,避免顶级准确率下降超过四倍。快速二次牛顿估计器仅需约9次Arnoldi迭代即可定位尼西莫里温度,较二分法提速六倍。所得集成模型将原1280维的MobileNetV2表示压缩至32维(ImageNet-10)或64维(ImageNet-100)。采用三图软集成,在ImageNet-10上达98.7%顶1准确率,在ImageNet-100上达84.92%。相较MobileNetV2,硬集成提升0.10%准确率,同时降低2.67倍浮点运算量;相比ResNet-50,软集成仅损失1.09%准确率,但浮点运算量减少29倍。创新点在于:(a) 建立图捕获集与代数拓扑缺陷的严格关联;(b) 提出高效尼西莫里温度估计方法;(c) 首次证明拓扑引导的LDPC图嵌入可用于高度压缩分类器。

原文摘要 · Abstract (English)

Modern multi-class image classification uses high-dimensional CNN features that incur large memory and computational costs and obscure the data manifold's geometry. Existing graph-based spectral classifiers work on synthetic or binary tasks but degrade on natural images with many classes because feature manifolds have non-trivial topology. We introduce a physics-inspired pipeline where frozen MobileNetV2 features are interpreted as Ising spins on a sparse multi-edge type quasi-cyclic LDPC graph, defining a Random-Bond Ising Model (RBIM). The model is operated at its Nishimori temperature -- where the smallest eigenvalue of the Bethe-Hessian matrix vanishes. A spectral-topological correspondence links trapping sets in the Tanner graph to topological invariants via poles of the Ihara-Bass zeta function, enabling systematic suppression of harmful substructures that otherwise reduce top-1 accuracy by more than a factor of four. A fast quadratic-Newton estimator finds the Nishimori temperature in $\sim 9$ Arnoldi iterations, a sixfold speed-up over bisection. The resulting ensembles compress the original $1280$-dimensional MobileNetV2 representation to $32$ dimensions (ImageNet-10) or $64$ dimensions (ImageNet-100). We achieve $98.7\%$ top-1 accuracy on ImageNet-10 and $84.92\%$ on ImageNet-100 using a three-graph soft ensemble. Relative to MobileNetV2, our hard ensemble increases accuracy by $0.10\%$ while reducing FLOPs by a factor of $2.67$. Against ResNet-50, the soft ensemble drops only 1.09% accuracy yet cuts FLOPs by $29\times$. The novelty lies in (a) establishing a rigorous link between graph trapping sets and algebraic-topological defects, (b) an efficient Nishimori-temperature estimator, and (c) demonstrating topology-guided LDPC graph embedding for highly compressed classifiers.

图像分类图神经网络压缩模型物理启发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。