arXiv:2603.14830cs.LGstat.ML2026-03中稿 · ICML

理论证明数据蒸馏可高效压缩非线性任务的低维信息。

Dataset Distillation Efficiently Encodes Low-Dimensional Representations from Gradient-Based Learning of Non-Linear Tasks

  • 基于梯度学习,将非线性任务的低维结构编码到合成数据中。
  • 仅需约 $\tildeΘ(r^2d + L)$ 的内存即可实现高泛化模型。
  • 适用于关注数据压缩效率与理论机制的研究者。

数据蒸馏是一种训练感知的数据压缩技术,近年来因其有效降低优化与存储成本而受到广泛关注。然而,其进展仍主要依赖经验。如何从训练过程中提取任务相关信息,并将其高效编码为合成数据点的机制仍不清晰。本文针对两层神经网络在宽度 $L$ 条件下的梯度训练,对数据蒸馏的实际算法进行理论分析。聚焦于称为多指标模型的非线性任务结构,证明该问题的低维特性可被高效编码至蒸馏数据中。所得数据在 $\tildeΘ(r^2d + L)$ 的内存复杂度下,即可重现具备高泛化能力的模型,其中 $d$ 和 $r$ 分别为任务的输入维度与内在维度。据我们所知,这是首个引入具体任务结构、利用内在维度量化压缩率,并仅通过梯度算法实现数据蒸馏的理论工作。

原文摘要 · Abstract (English)

Dataset distillation, a training-aware data compression technique, has recently attracted increasing attention as an effective tool for mitigating costs of optimization and data storage. However, progress remains largely empirical. Mechanisms underlying the extraction of task-relevant information from the training process and the efficient encoding of such information into synthetic data points remain elusive. In this paper, we theoretically analyze practical algorithms of dataset distillation applied to the gradient-based training of two-layer neural networks with width $L$. By focusing on a non-linear task structure called multi-index model, we prove that the low-dimensional structure of the problem is efficiently encoded into the resulting distilled data. This dataset reproduces a model with high generalization ability for a required memory complexity of $\tildeΘ$$(r^2d+L)$, where $d$ and $r$ are the input and intrinsic dimensions of the task. To the best of our knowledge, this is one of the first theoretical works that include a specific task structure, leverage its intrinsic dimensionality to quantify the compression rate and study dataset distillation implemented solely via gradient-based algorithms.

数据蒸馏理论分析低维编码神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。