arXiv:2505.15267cs.CV2025-05被引 1

用对比学习提升小数据集蒸馏效果,让合成数据更丰富清晰。

Contrastive Learning-Enhanced Trajectory Matching for Small-Scale Dataset Distillation

  • 在图像生成中引入对比学习,增强样本特征区分度
  • 极小数据量下仍能保持高视觉质量和语义信息
  • 适合边缘设备部署和快速原型开发场景

在资源受限环境(如边缘设备或快速原型设计)中部署机器学习模型,亟需将大规模真实数据集蒸馏为规模更小但信息丰富的合成数据集。现有数据集蒸馏方法,尤其是轨迹匹配类方法,通过优化合成数据使模型在合成数据上的训练轨迹与真实数据一致。尽管在中等规模合成数据上表现良好,但在极端样本稀缺情况下难以保留语义丰富性。为此,我们提出一种新方法,在图像生成过程中融合对比学习,通过显式最大化实例级特征区分度,生成更具信息量和多样性的合成样本,即使在数据规模严重受限时亦然。实验表明,该方法显著提升了在极小规模合成数据上训练模型的性能,不仅引导更有效的特征表示,还大幅提高合成图像的视觉保真度。结果证明,本方法在极小数据场景下优于现有蒸馏技术。

原文摘要 · Abstract (English)

Deploying machine learning models in resource-constrained environments, such as edge devices or rapid prototyping scenarios, increasingly demands distillation of large datasets into significantly smaller yet informative synthetic datasets. Current dataset distillation techniques, particularly Trajectory Matching methods, optimize synthetic data so that the model's training trajectory on synthetic samples mirrors that on real data. While demonstrating efficacy on medium-scale synthetic datasets, these methods fail to adequately preserve semantic richness under extreme sample scarcity. To address this limitation, we propose a novel dataset distillation method integrating contrastive learning during image synthesis. By explicitly maximizing instance-level feature discrimination, our approach produces more informative and diverse synthetic samples, even when dataset sizes are significantly constrained. Experimental results demonstrate that incorporating contrastive learning substantially enhances the performance of models trained on very small-scale synthetic datasets. This integration not only guides more effective feature representation but also significantly improves the visual fidelity of the synthesized images. Experimental results demonstrate that our method achieves notable performance improvements over existing distillation techniques, especially in scenarios with extremely limited synthetic data.

数据蒸馏对比学习小样本边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。