arXiv:2508.19659cs.LG2025-08

提出SCAR框架,用四维指标刻画多模态数据内在结构。

SCAR: A Characterization Scheme for Multi-Modal Dataset

  • 从规模、覆盖、真实性和丰富性四维度量化数据结构特性。
  • 发现数据扩展中保持不变的特征,可预测模型泛化能力。
  • 适用于高效构建最小有效数据集,适合数据优化研究者。

基础模型在多种任务中表现出卓越泛化能力,主要依赖于训练数据的特性。现有数据驱动方法如剪枝与压缩虽能优化训练过程,但缺乏对数据属性如何影响泛化的理论理解,尤其在样本扩展场景下。传统视角过度关注数据量和训练效率,常忽视数据质量的结构性特征。为此,本文提出SCAR——一种基于规模(Scale)、覆盖(Coverage)、真实性(Authenticity)和丰富性(Richness)四个维度的系统性数据表征方案。该方案捕捉在数据缩放下仍保持稳定的内在结构特征,为数据理解提供稳健通用的基础。基于此,我们定义了「Foundation Data」:一个最小数据子集,能保留完整数据集的泛化行为而无需模型重训练。通过将单模态任务建模为阶跃函数,估计基础数据规模分布,以捕获跨模态的阶跃式泛化偏差。最后,设计了基于此偏差的SCAR引导数据补全策略,实现多模态数据中特定模态特性的高效、感知模态的扩充。在多种多模态数据集和模型架构上的实验验证了SCAR在预测数据效用和指导数据获取方面的有效性。代码已开源。

原文摘要 · Abstract (English)

Foundation models exhibit remarkable generalization across diverse tasks, largely driven by the characteristics of their training data. Recent data-centric methods like pruning and compression aim to optimize training but offer limited theoretical insight into how data properties affect generalization, especially the data characteristics in sample scaling. Traditional perspectives further constrain progress by focusing predominantly on data quantity and training efficiency, often overlooking structural aspects of data quality. In this study, we introduce SCAR, a principled scheme for characterizing the intrinsic structural properties of datasets across four key measures: Scale, Coverage, Authenticity, and Richness. Unlike prior data-centric measures, SCAR captures stable characteristics that remain invariant under dataset scaling, providing a robust and general foundation for data understanding. Leveraging these structural properties, we introduce Foundation Data-a minimal subset that preserves the generalization behavior of the full dataset without requiring model-specific retraining. We model single-modality tasks as step functions and estimate the distribution of the foundation data size to capture step-wise generalization bias across modalities in the target multi-modal dataset. Finally, we develop a SCAR-guided data completion strategy based on this generalization bias, which enables efficient, modality-aware expansion of modality-specific characteristics in multimodal datasets. Experiments across diverse multi-modal datasets and model architectures validate the effectiveness of SCAR in predicting data utility and guiding data acquisition. Code is available at https://github.com/McAloma/SCAR.

多模态数据表征基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。