提出一种统一方法,实现多模态数据压缩并保持性能。
Omnimodal Dataset Distillation via High-order Proxy Alignment

- 用紧凑代理捕捉高阶跨模态对齐,避免成对建模复杂度。
- 在多个基准上实现优于现有方法的压缩与性能平衡。
- 适合需要高效处理多源异构数据的场景,如多模态学习。
数据蒸馏可将大规模数据集压缩为紧凑的合成数据集,同时保持训练性能,但现有方法多局限于单模态或双模态场景。将数据蒸馏扩展至三模态及以上(即全模态数据蒸馏)仍缺乏探索且极具挑战,因模态间异质性增强及跨模态交互复杂。本文识别出全模态设置下终点差异的关键决定因素,其随模态数量增加而加剧。为此,提出HoPA方法,通过紧凑代理捕获高阶跨模态对齐,兼容轨迹匹配。通过共享相似性结构抽象全模态对齐,避免了成对模态建模的组合爆炸,实现了异构模态间的可扩展联合蒸馏。从谱理论视角的分析验证了该方法相较于双模态蒸馏技术的合理性。大量实验证明,所提方法在多种基准上均显著优于现有竞争方法,实现了更优的压缩-性能权衡。源代码将公开发布。
原文摘要 · Abstract (English)
Dataset distillation compresses large-scale datasets into compact synthetic sets while preserving training performance, but existing methods are largely restricted to single-modal or bimodal settings. Extending dataset distillation to scenarios involving more than two modalities, i.e., Omnimodal Dataset Distillation, remains underexplored and challenging due to increased heterogeneity and complex cross-modal interactions. In this work, we identify the key determinant that bounds the endpoint discrepancy in the omnimodal setting, which is exacerbated with an increasing number of modalities. To this end, we propose HoPA, a unified method that captures high-order cross-modal alignments via a compact proxy, which is compatible with trajectory matching as well. By abstracting omnimodal alignment with a shared similarity structure, our method avoids the combinatorial complexity of pairwise modality modeling and enables scalable joint distillation across heterogeneous modalities. Theoretical analysis from the spectral perspective reveals the rationality of our proposed method against bimodal dataset distillation techniques. Extensive experiments on various benchmarks demonstrate that the proposed method achieves superior compression-performance trade-offs compared to existing competitors. The source code will be publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。