arXiv:2601.21296cs.LGcs.AI2026-01中稿 · ICLR被引 3

提出新框架提升数据蒸馏的有用性和信息量,效果显著优于旧方法。

Grounding and Enhancing Informativeness and Utility in Dataset Distillation

  • 用博弈论和贡献度分析提取关键信息,提升样本价值
  • 在ImageNet-1K上比当前最佳方法高6.1%准确率
  • 适合关注高效训练与小规模数据集设计的研究者

数据蒸馏旨在从大规模真实数据集中生成紧凑的数据集。现有方法多依赖启发式策略平衡效率与质量,但原始数据与合成数据间的关系仍缺乏深入探索。本文在理论框架下重新审视基于知识蒸馏的数据蒸馏方法,引入‘信息量’与‘效用’两个概念,分别刻画样本中的关键信息和训练集中重要样本。基于此,我们数学定义了最优数据蒸馏。进而提出InfoUtil框架,通过两部分实现平衡:(1) 基于谢尔平利值(Shapley Value)的博弈论信息量最大化,以提取样本关键信息;(2) 基于梯度范数的全局影响力样本选择,实现效用优化。实验表明,在ImageNet-1K数据集上使用ResNet-18时,该方法相比先前最先进方法提升6.1%性能。

原文摘要 · Abstract (English)

Dataset Distillation (DD) seeks to create a compact dataset from a large, real-world dataset. While recent methods often rely on heuristic approaches to balance efficiency and quality, the fundamental relationship between original and synthetic data remains underexplored. This paper revisits knowledge distillation-based dataset distillation within a solid theoretical framework. We introduce the concepts of Informativeness and Utility, capturing crucial information within a sample and essential samples in the training set, respectively. Building on these principles, we define optimal dataset distillation mathematically. We then present InfoUtil, a framework that balances informativeness and utility in synthesizing the distilled dataset. InfoUtil incorporates two key components: (1) game-theoretic informativeness maximization using Shapley Value attribution to extract key information from samples, and (2) principled utility maximization by selecting globally influential samples based on Gradient Norm. These components ensure that the distilled dataset is both informative and utility-optimized. Experiments demonstrate that our method achieves a 6.1\% performance improvement over the previous state-of-the-art approach on ImageNet-1K dataset using ResNet-18.

数据蒸馏信息量效用优化模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。