用语义筛选与增强,让工业级数据训练更高效、更精准。
SSE: Multimodal Semantic Data Selection and Enrichment for Industrial-scale Data Assimilation
- 基于语义多样性筛选关键数据,再从海量无标签数据中挖掘新内容。
- 用更小的数据集维持甚至提升模型性能,不超原始数据规模。
- 借助基础模型解释每条数据语义,适合工业场景的数据优化需求。
近年来,人工智能收集的数据量已达到难以管理的程度。在自动驾驶等工业应用中,模型训练计算预算被突破,性能却趋于饱和,而数据仍在持续涌入。为应对数据洪流,我们提出一种框架,用于选择最具语义多样性和重要性的数据子集,并通过从大规模无标签数据池中发现有意义的新数据来进一步语义增强。关键在于,我们利用基础模型为每个数据点生成语义表示,实现可解释性。定量结果表明,我们的语义选择与增强框架(SSE)能够:a)在更小的训练数据集上保持模型性能;b)在不超出原数据集大小的前提下,通过增强小数据集提升模型性能。因此,我们证明了语义多样性对最优数据选择和模型表现至关重要。
原文摘要 · Abstract (English)
In recent years, the data collected for artificial intelligence has grown to an unmanageable amount. Particularly within industrial applications, such as autonomous vehicles, model training computation budgets are being exceeded while model performance is saturating -- and yet more data continues to pour in. To navigate the flood of data, we propose a framework to select the most semantically diverse and important dataset portion. Then, we further semantically enrich it by discovering meaningful new data from a massive unlabeled data pool. Importantly, we can provide explainability by leveraging foundation models to generate semantics for every data point. We quantitatively show that our Semantic Selection and Enrichment framework (SSE) can a) successfully maintain model performance with a smaller training dataset and b) improve model performance by enriching the smaller dataset without exceeding the original dataset size. Consequently, we demonstrate that semantic diversity is imperative for optimal data selection and model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。