arXiv:2512.13083cs.CVcs.LG2025-12中稿 · WACV 2026

提升数据浓缩生成数据集的多样性,减少冗余。

DiRe: Diversity-promoting Regularization for Dataset Condensation

  • 引入余弦相似度与欧氏距离结合的多样性正则化方法
  • 在CIFAR-10到ImageNet-1K上显著提升泛化与多样性表现
  • 可直接用于主流浓缩方法,无需改动框架

数据浓缩的目标是生成一个小型数据集,以复现大型原始数据集的训练效果。现有浓缩方法生成的数据集存在明显冗余,亟需降低冗余并提升多样性。为此,我们提出一种直观的多样性正则化器(DiRe),由余弦相似度与欧氏距离构成,可直接应用于多种先进浓缩方法。大量实验表明,加入该正则化器后,主流浓缩方法在从CIFAR-10到ImageNet-1K的多个基准数据集上,均在泛化性和多样性指标上获得提升。

原文摘要 · Abstract (English)

In Dataset Condensation, the goal is to synthesize a small dataset that replicates the training utility of a large original dataset. Existing condensation methods synthesize datasets with significant redundancy, so there is a dire need to reduce redundancy and improve the diversity of the synthesized datasets. To tackle this, we propose an intuitive Diversity Regularizer (DiRe) composed of cosine similarity and Euclidean distance, which can be applied off-the-shelf to various state-of-the-art condensation methods. Through extensive experiments, we demonstrate that the addition of our regularizer improves state-of-the-art condensation methods on various benchmark datasets from CIFAR-10 to ImageNet-1K with respect to generalization and diversity metrics.

数据浓缩多样性正则化图像分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。