arXiv:2509.19743cs.CV2025-09中稿 · ICLR被引 2

提出标准化评估方法,揭示数据蒸馏性能差异多因评测不一致。

Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive Evaluation

  • 分离教师模型训练与合成数据生成,降低计算开销。
  • 发现现有方法性能差异主要源于评估流程不统一。
  • 建立统一基准,推动公平可复现的数据蒸馏研究。

数据蒸馏旨在生成紧凑的合成数据集,使模型在这些数据上训练后性能接近在完整真实数据集上训练的结果,同时大幅降低存储与计算成本。早期双层优化方法(如MTT)在小规模数据集上表现良好,但因计算开销高而难以扩展。近期解耦式蒸馏方法(如SRe²L)将教师模型预训练与合成数据生成分离,并在后评估阶段引入随机数据增强和逐轮软标签以提升性能与泛化能力。然而,现有解耦方法存在评估协议不一致的问题,阻碍了领域进展。本文提出修正解耦数据蒸馏(RD³),系统研究不同后评估设置对测试准确率的影响,进一步检验已有方法报告的性能差异是否反映真正的方法进步,还是源于评估流程差异。分析表明,大部分性能波动源于评估不一致而非合成数据本身质量差异。我们还识别出跨场景提升蒸馏效果的通用策略。通过建立标准化基准与严格评估协议,RD³为未来数据蒸馏研究提供了公平、可复现的比较基础。

原文摘要 · Abstract (English)

Dataset distillation aims to generate compact synthetic datasets that enable models trained on them to achieve performance comparable to those trained on full real datasets, while substantially reducing storage and computational costs. Early bi-level optimization methods (e.g., MTT) have shown promising results on small-scale datasets, but their scalability is limited by high computational overhead. To address this limitation, recent decoupled dataset distillation methods (e.g., SRe$^2$L) separate the teacher model pre-training from the synthetic data generation process. These methods also introduce random data augmentation and epoch-wise soft labels during the post-evaluation phase to improve performance and generalization. However, existing decoupled distillation methods suffer from inconsistent post-evaluation protocols, which hinders progress in the field. In this work, we propose Rectified Decoupled Dataset Distillation (RD$^3$), and systematically investigate how different post-evaluation settings affect test accuracy. We further examine whether the reported performance differences across existing methods reflect true methodological advances or stem from discrepancies in evaluation procedures. Our analysis reveals that much of the performance variation can be attributed to inconsistent evaluation rather than differences in the intrinsic quality of the synthetic data. In addition, we identify general strategies that improve the effectiveness of distilled datasets across settings. By establishing a standardized benchmark and rigorous evaluation protocol, RD$^3$ provides a foundation for fair and reproducible comparisons in future dataset distillation research.

数据蒸馏评估标准模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。