arXiv:2501.04202cs.CVcs.AI2025-01中稿 · ICASSP 2025被引 5

用自知识蒸馏生成更精准的压缩数据集,提升模型训练效率。

Generative Dataset Distillation Based on Self-knowledge Distillation

  • 引入自知识蒸馏机制,优化合成数据与原始数据的分布匹配。
  • 对预测逻辑值进行标准化处理,提升分布对齐精度。
  • 适用于需要高效训练的场景,尤其适合资源受限环境。

数据集蒸馏是一种有效降低模型训练成本和复杂度的技术,通过将大规模数据集压缩为更小、更高效的版本来保持性能。本文提出一种新型生成式数据集蒸馏方法,可提升预测输出逻辑值对齐的准确性。该方法结合自知识蒸馏,实现合成数据与原始数据间更精确的分布匹配,从而捕捉数据的整体结构与内在关系。为进一步提升对齐精度,我们在分布匹配前对逻辑值引入标准化步骤,确保逻辑值范围一致。通过大量实验验证,本方法优于现有最先进方法,在数据蒸馏性能上表现更优。

原文摘要 · Abstract (English)

Dataset distillation is an effective technique for reducing the cost and complexity of model training while maintaining performance by compressing large datasets into smaller, more efficient versions. In this paper, we present a novel generative dataset distillation method that can improve the accuracy of aligning prediction logits. Our approach integrates self-knowledge distillation to achieve more precise distribution matching between the synthetic and original data, thereby capturing the overall structure and relationships within the data. To further improve the accuracy of alignment, we introduce a standardization step on the logits before performing distribution matching, ensuring consistency in the range of logits. Through extensive experiments, we demonstrate that our method outperforms existing state-of-the-art methods, resulting in superior distillation performance.

数据蒸馏知识蒸馏生成模型高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。