arXiv:2509.10367cs.LG2025-09被引 1

用差异度量统一定义数据压缩,提升模型泛化与鲁棒性。

A Discrepancy-Based Perspective on Dataset Condensation

  • 基于差异度量构建统一的数据压缩框架
  • 压缩后模型性能可媲美甚至超越原数据训练结果
  • 支持鲁棒性、隐私等多重目标,适用范围更广

给定一个包含有限元素的原始数据集 $/mathcal{T} = \{\mathbf{x}_i\}_{i = 1}^N$,数据集压缩(DC)的目标是构造一个显著更小的合成数据集 $/mathcal{S} = \{\tilde{\mathbf{x}}_j\}_{j = 1}^M$($M \ll N$),使得从零开始在 $/mathcal{S}$ 上训练的模型能达到与在 $/mathcal{T}$ 上训练相当甚至更优的泛化性能。近期研究揭示了数据压缩与用少量点逼近原始数据分布之间的紧密联系。本文提出一个统一框架,涵盖现有 DC 方法,并将任务特定的压缩概念形式化为基于差异度量的通用定义,用于量化不同场景下概率分布间的距离。该框架将 DC 目标扩展至泛化之外,支持鲁棒性、隐私等额外目标。

原文摘要 · Abstract (English)

Given a dataset of finitely many elements $\mathcal{T} = \{\mathbf{x}_i\}_{i = 1}^N$, the goal of dataset condensation (DC) is to construct a synthetic dataset $\mathcal{S} = \{\tilde{\mathbf{x}}_j\}_{j = 1}^M$ which is significantly smaller ($M \ll N$) such that a model trained from scratch on $\mathcal{S}$ achieves comparable or even superior generalization performance to a model trained on $\mathcal{T}$. Recent advances in DC reveal a close connection to the problem of approximating the data distribution represented by $\mathcal{T}$ with a reduced set of points. In this work, we present a unified framework that encompasses existing DC methods and extend the task-specific notion of DC to a more general and formal definition using notions of discrepancy, which quantify the distance between probability distribution in different regimes. Our framework broadens the objective of DC beyond generalization, accommodating additional objectives such as robustness, privacy, and other desirable properties.

数据压缩分布逼近泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。