arXiv:2603.02411cs.CVcs.AI2026-03中稿 · CVPR

提出联合优化数据量与精度的压缩框架,提升单位比特效率

From Fewer Samples to Fewer Bits: Reframing Dataset Distillation as Joint Optimization of Precision and Compactness

  • 在蒸馏过程中引入可微量化模块,同步优化合成样本与量化参数
  • 在图像分类和3GPP波束管理任务中,单位比特准确率超越现有方法
  • 支持自适应量化,更好保留信息密集区域的细节,适合资源受限场景

数据蒸馏(DD)将大规模数据集压缩为紧凑的合成数据以保持训练性能。然而,现有方法主要关注样本数量减少,对数据精度及其效率影响考虑不足。本文提出量化感知数据蒸馏(QuADD),在固定比特预算下联合优化数据紧凑性与精度。QuADD 在蒸馏循环中集成可微量化模块,实现合成样本与量化参数的端到端协同优化。基于率失真视角,我们实证分析了样本数量与精度间比特分配对学习性能的影响。该框架支持均匀与非均匀自适应量化,后者可从数据中学习量化层级,更优表示信息密集区域。在图像分类与3GPP波束管理任务上的实验表明,QuADD 在每比特准确率上超越现有DD及后量化基线,确立了信息高效数据蒸馏的新标准。

原文摘要 · Abstract (English)

Dataset Distillation (DD) compresses large datasets into compact synthetic ones that maintain training performance. However, current methods mainly target sample reduction, with limited consideration of data precision and its impact on efficiency. We propose Quantization-aware Dataset Distillation (QuADD), a unified framework that jointly optimizes dataset compactness and precision under fixed bit budgets. QuADD integrates a differentiable quantization module within the distillation loop, enabling end-to-end co-optimization of synthetic samples and quantization parameters. Guided by the rate-distortion perspective, we empirically analyze how bit allocation between sample count and precision influences learning performance. Our framework supports both uniform and adaptive non-uniform quantization, where the latter learns quantization levels from data to represent information-dense regions better. Experiments on image classification and 3GPP beam management tasks show that QuADD surpasses existing DD and post-quantized baselines in accuracy per bit, establishing a new standard for information-efficient dataset distillation.

数据蒸馏量化压缩高效学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。