用对偶间隙扩展核诱导点,让数据蒸馏支持更多损失函数。
Generalized Kernel Inducing Points by Duality Gap for Dataset Distillation
- 基于对偶间隙构造新诱导点,避免双层优化
- 在MNIST和CIFAR-10上保持高效且支持交叉熵等损失
- 理论证明测试误差与预测一致性有上界,适合分类任务
我们提出对偶间隙诱导点(DGKIP),作为核诱导点(KIP)方法在数据蒸馏中的扩展。现有数据蒸馏方法常依赖双层优化,而DGKIP通过凸规划的对偶理论消除了这一需求。尽管KIP方法已避免双层优化,但仅适用于平方损失,不支持交叉熵或合页损失等更适用于分类任务的损失函数。DGKIP通过利用数据蒸馏后参数变化的对偶间隙上界,实现了对多种损失函数的适用性。我们还通过理论分析给出了蒸馏后测试误差和预测一致性的上界。在标准基准如MNIST和CIFAR-10上的实验表明,DGKIP在保持KIP效率的同时,具备更广的适用性和稳健性能。
原文摘要 · Abstract (English)
We propose Duality Gap KIP (DGKIP), an extension of the Kernel Inducing Points (KIP) method for dataset distillation. While existing dataset distillation methods often rely on bi-level optimization, DGKIP eliminates the need for such optimization by leveraging duality theory in convex programming. The KIP method has been introduced as a way to avoid bi-level optimization; however, it is limited to the squared loss and does not support other loss functions (e.g., cross-entropy or hinge loss) that are more suitable for classification tasks. DGKIP addresses this limitation by exploiting an upper bound on parameter changes after dataset distillation using the duality gap, enabling its application to a wider range of loss functions. We also characterize theoretical properties of DGKIP by providing upper bounds on the test error and prediction consistency after dataset distillation. Experimental results on standard benchmarks such as MNIST and CIFAR-10 demonstrate that DGKIP retains the efficiency of KIP while offering broader applicability and robust performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。