arXiv:2511.08263cs.CVcs.AI2025-11AAAI被引 4

用ImageBind统一特征空间压缩多模态数据,仅5个样本就实现全数据集效果。

ImagebindDC: Compressing Multi-modal Data with Imagebind-based Condensation

  • 基于ImageBind的统一特征空间,用特征函数损失实现精确统计对齐。
  • 每类仅需5个合成样本,NYU-v2上达到与全量数据相当的无损性能。
  • 适合需要高效多模态数据压缩的研究者,尤其关注训练加速与资源节省。

数据凝缩技术旨在从大规模数据集中合成紧凑数据集以实现高效模型训练,但在多模态场景中常因难以保持复杂的跨模态依赖而失效。为此,我们提出ImageBindDC,一种在ImageBind统一特征空间中的新型数据凝缩框架。该方法超越传统分布匹配,采用强大的特征函数(CF)损失,在傅里叶域实现精确的无穷阶矩匹配,促进更精准的统计对齐。我们的目标强制实现三个关键层面的分布一致性:(i) 单模态对齐,匹配各模态内合成与真实数据的统计特性;(ii) 跨模态对齐,通过匹配混合真实-合成数据对的分布来保留成对语义;(iii) 联合模态对齐,通过对齐真实数据对与其合成对应物的联合分布,捕捉完整的多变量数据结构。大量实验表明ImageBindDC的有效性:在NYU-v2数据集上,每类仅使用5个凝缩样本即可达到与完整数据集训练相当的无损性能,相比此前最优方法提升8.2%绝对准确率,且凝缩时间减少4倍以上。

原文摘要 · Abstract (English)

Data condensation techniques aim to synthesize a compact dataset from a larger one to enable efficient model training, yet while successful in unimodal settings, they often fail in multimodal scenarios where preserving intricate inter-modal dependencies is crucial. To address this, we introduce ImageBindDC, a novel data condensation framework operating within the unified feature space of ImageBind. Our approach moves beyond conventional distribution-matching by employing a powerful Characteristic Function (CF) loss, which operates in the Fourier domain to facilitate a more precise statistical alignment via exact infinite moment matching. We design our objective to enforce three critical levels of distributional consistency: (i) uni-modal alignment, which matches the statistical properties of synthetic and real data within each modality; (ii) cross-modal alignment, which preserves pairwise semantics by matching the distributions of hybrid real-synthetic data pairs; and (iii) joint-modal alignment, which captures the complete multivariate data structure by aligning the joint distribution of real data pairs with their synthetic counterparts. Extensive experiments highlight the effectiveness of ImageBindDC: on the NYU-v2 dataset, a model trained on just 5 condensed datapoints per class achieves lossless performance comparable to one trained on the full dataset, achieving a new state-of-the-art with an 8.2\% absolute improvement over the previous best method and more than 4$\times$ less condensation time.

多模态数据压缩图像生成高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。