通过分析离散视觉标记的结构,提升数据蒸馏效果。
Structural Assessment for Understanding and Guiding Dataset Distillation in Discrete Token Space

- 用离散标记统计分析数据组成结构,定义结构得分。
- 平衡的标记组合能显著提升验证性能。
- 适合关注数据蒸馏机制与生成式蒸馏的研究者。
数据蒸馏(DD)已证明可在保持准确率的同时降低训练成本。然而,为何某些蒸馏数据集比其他更有效仍不明确。本文从离散视觉标记器的角度探究此问题。以往工作多强调全局数据分布匹配,我们提出有效性取决于捕捉到的语义概念及其组合方式。离散视觉标记器提供有限词汇,可直接分析组合结构。通过标记级统计量化分析,我们引入结构得分以衡量标记组合的充分性。发现标记组合均衡的蒸馏数据集具有更高验证性能。此外,与原始数据的偏离并不必然损害性能。我们进一步表明,离散标记空间中结构得分高的样本可有效引导基于扩散的蒸馏。研究凸显了标记组合在数据蒸馏中的重要性,为分布相似性之外提供了原则性补充。
原文摘要 · Abstract (English)
Dataset distillation (DD) has proven to reduce training cost while preserving accuracy. While promising, the factors that make one distilled dataset more effective than another remain poorly understood. In this work, we investigate this question through the lens of discrete visual tokenizers. Whereas many prior DD efforts emphasize matching global data distributions, we suggest that the effectiveness depends on which semantic concepts are captured and how they are composed. Discrete visual tokenizers provide a finite vocabulary that enables direct statistical analysis of such compositional structure. Through quantitative analysis of token-level statistics, we introduce the structural score to measure the adequacy of token compositions. We observe that distilled datasets with balanced token composition yield higher validation performance. On the other hand, divergence from the original data does not necessarily harm performance. We further show that samples with high structural scores in the discrete token space can effectively guide diffusion-based DD. Our findings highlight the importance of token composition in dataset effectiveness, offering a principled complement to distributional similarity considerations in DD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。