针对多模态数据集选择中的语义失衡与分布偏差问题,提出一种融合多尺度拓扑结构的高效选子集方法。
CAST: Collapse-Aware multi-Scale Topology Fusion for Multimodal Coreset Selection

- 构建图像与文本模态拓扑,通过跨模态融合与局部坍缩感知优化统一拓扑结构
- 在扩散小波域实现多尺度分布匹配,提升子集对原数据全局与局部结构的逼近能力
- 引入局部软关系覆盖机制,有效避免密集区域重复采样,增强冗余感知能力
大型多模态模型训练依赖海量图文数据集,带来巨大计算开销。数据集选择通过识别高信息量子集提供可行方案,但现有方法存在两大缺陷:(i) 单模态主导的采样方式忽略多模态数据中细粒度的跨模态信息不平衡,导致另一模态语义损失;(ii) 粗粒度评分驱动的采样方式使子集偏向评分模型,难以保证子集与原始数据的分布一致性。同时,现有分布匹配与离散采样策略常无法兼顾全局语义结构、局部细粒度细节及密集区域的冗余感知覆盖。为此,我们提出CAST框架——一种面向多模态子集选择的崩溃感知多尺度拓扑融合方法。首先构建图像与文本模态拓扑,通过局部坍缩感知修正与跨模态融合生成统一拓扑;其次在扩散小波域引入多尺度分布匹配准则,促使子集在多尺度上逼近原数据分布;最后设计局部软关系覆盖机制,将纯几何覆盖扩展为关系感知的间接覆盖,抑制密集簇中的冗余采样。在Flickr30K与MS-COCO上的大量实验表明,CAST优于现有基线方法,在跨架构泛化性和能效方面显著领先于当前最优多模态合成方法。
原文摘要 · Abstract (English)
The training of large multimodal models fundamentally relies on massive image-text datasets, which inevitably incur prohibitive computational overhead. Dataset selection offers a promising paradigm by identifying a highly informative coreset. However, existing approaches suffer from two critical limitations: (i) single-modality-dominated sampling methods, which ignore the fine-grained cross-modal information imbalance inherent in multimodal datasets and thus lead to semantic loss in the other modality; and (ii) coarse-grained sample-scoring-based sampling methods, where the selected coreset tends to be biased toward the scoring model, making it difficult to guarantee distributional equivalence between the coreset and the original dataset. Meanwhile, existing distribution matching and discrete sampling strategies often fail to jointly account for global semantic structure, local fine-grained details, and redundancy-aware coverage in dense regions. To this end, we propose CAST, a Collapse-Aware multi-Scale Topology fusion framework for multimodal coreset selection. We first construct image- and text-modality topologies, and derive a unified topology via local-collapse-aware refinement and cross-modal fusion. We then introduce a multi-scale distribution matching criterion in the diffusion wavelet domain, encouraging the coreset to approximate the original dataset at multiple scales. Finally, we introduce a local soft relational coverage mechanism that extends pure geometric coverage to relation-aware indirect coverage, penalizing redundant selections in dense clusters. Extensive experiments on Flickr30K and MS-COCO show that CAST outperforms existing dataset selection baselines, showcasing great superiority in cross-architecture generalization and energy efficiency over state-of-the-art multimodal synthesis methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。