arXiv:2602.11320cs.LG2026-02被引 1

用数据压缩降低神经正切核计算量,提速近百倍。

Efficient Analysis of the Distilled Neural Tangent Kernel

  • 用神经正切核优化的数据蒸馏压缩输入维度
  • 雅可比矩阵计算量减少20-100倍,有效秩保持不变
  • 适合需要高效核方法的机器学习研究者

神经正切核(NTK)方法受限于在大量数据点上计算大型雅可比矩阵。现有方法主要通过投影和随机采样降低开销。本文提出利用NTK调优的数据蒸馏压缩数据维度本身,证明输入数据张成的神经正切空间可通过蒸馏数据重构,使所需雅可比计算量减少20至100倍。进一步发现,类内NTK矩阵具有低有效秩,且该特性在压缩后得以保留。基于此,我们提出蒸馏神经正切核(DNTK),结合NTK调优的数据蒸馏与先进投影方法,将NTK计算复杂度降低达五数量级,同时保持核结构和预测性能。

原文摘要 · Abstract (English)

Neural tangent kernel (NTK) methods are computationally limited by the need to evaluate large Jacobians across many data points. Existing approaches reduce this cost primarily through projecting and sketching the Jacobian. We show that NTK computation can also be reduced by compressing the data dimension itself using NTK-tuned dataset distillation. We demonstrate that the neural tangent space spanned by the input data can be induced by dataset distillation, yielding a 20-100$\times$ reduction in required Jacobian calculations. We further show that per-class NTK matrices have low effective rank that is preserved by this reduction. Building on these insights, we propose the distilled neural tangent kernel (DNTK), which combines NTK-tuned dataset distillation with state-of-the-art projection methods to reduce up NTK computational complexity by up to five orders of magnitude while preserving kernel structure and predictive performance.

神经正切核数据蒸馏计算效率降维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。