arXiv:2510.15162cs.CVcs.CL2025-10EMNLP被引 1

用合成数据训练统一模型,自动筛选图文数据质量。

Train a Unified Multimodal Data Quality Classifier with Synthetic Data

  • 用合成图文对构建多级标签数据,训练统一质检模型。
  • 在DataComp和OBELICS数据集上过滤后,模型零样本能力显著提升。
  • 适合做多模态预训练数据清洗的研究者与工程师使用。

多模态大语言模型(MLLM)持续在图像-文本标题数据与交错文档数据的混合数据上进行预训练,但针对图像-文本交错文档数据的高质量数据过滤仍处于探索阶段。本文提出一种高效MLLM作为统一多模态数据质量分类器(UniFilter),用于过滤高质量图像-文本标题数据和交错数据。为解决多样标注数据收集难题,我们引入半合成方法:利用现成原始图像,生成四个质量等级的对应文本,高效构建用于训练UniFilter的样本-评分对。将UniFilter应用于DataComp标题数据集和OBELICS图像-文本交错数据集,筛选出高质量数据。基于过滤后数据预训练的MLLM在零样本推理和上下文学习能力上显著优于基线模型。经视觉监督微调后,这些由UniFilter引导的MLLM在多个基准测试中表现更优,凸显高质量多模态预训练的下游价值。我们公开了训练UniFilter所用合成数据、模型检查点及经过UniFilter筛选的OBELICS-HQ高质量交错文档子集,供社区复现与进一步开发。

原文摘要 · Abstract (English)

The Multimodal Large Language Models (MLLMs) are continually pre-trained on a mixture of image-text caption data and interleaved document data, while the high-quality data filtering towards image-text interleaved document data is under-explored. We propose to train an efficient MLLM as a Unified Mulitmodal Data Quality Classifier to Filter both high-quality image-text caption and interleaved data (UniFilter). To address the challenge of collecting diverse labeled multimodal data, we introduce a semi-synthetic approach that leverages readily available raw images and generates corresponding text across four quality levels. This method enables efficient creation of sample-score pairs for both caption and interleaved document data to train UniFilter. We apply UniFilter to curate high-quality caption data from DataComp caption dataset and interleaved data from the OBELICS image-text interleaved dataset. MLLMs pre-trained on the filtered data demonstrate significantly enhanced capabilities compared to those trained on baseline-filtered data, achieving stronger zero-shot reasoning and in-context learning capabilities. After visual supervised fine-tuning, these UniFilter-induced MLLMs achieve stronger performance on various benchmarks, highlighting the downstream benefits of high-quality multimodal pre-training. We release the synthetic training data used for training UniFilter, the UniFilter model checkpoints, and the high-quality interleaved document subset OBELICS-HQ, curated by UniFilter, to the community for reproduction and further development.

多模态数据清洗合成数据MLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。