构建首个保留文档结构的阿拉伯语多模态数据集
Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora
- 基于Common Crawl构建阿拉伯语多模态数据,输出markdown格式
- 相比纯文本数据集,保留图文交错结构与文档层级
- 适合多模态预训练与阿拉伯语语言模型研究
大语言模型(LLMs)和大视觉语言模型(LMMs)的性能高度依赖于预训练数据的质量与规模。最新研究表明,在自然文档中以图文交错形式训练的大规模多模态模型,在多个基准测试中表现优于仅使用图像-文本对训练的模型,其优势源于先进预训练模型对语义对齐、图像序列一致性及文本连贯性的强化。然而,针对阿拉伯语,缺乏高质量、保留文档结构的多模态数据集,限制了相关进展。本文提出Wasm数据处理流水线,对Common Crawl数据集进行处理,构建首个提供markdown输出的阿拉伯语多模态数据集。与现有仅聚焦文本提取的阿拉伯语语料库不同,该方法在保持网页内容结构完整性的同时,兼顾纯文本与多模态预训练的灵活性。我们对本方案与主流数据集处理流程进行了全面对比分析,揭示过滤策略的共性,并论证设计选择的合理性。为支持后续研究,我们公开发布代表性数据样本及完整的多模态处理管道。
原文摘要 · Abstract (English)
The performance of large language models (LLMs) and large multimodal models (LMMs) depends heavily on the quality and scale of their pre-training datasets. Recent research shows that large multimodal models trained on natural documents where images and text are interleaved outperform those trained only on image-text pairs across a wide range of benchmarks, leveraging advanced pre-trained models to enforce semantic alignment, image-sequence consistency, and textual coherence. For Arabic, however, the lack of high-quality multimodal datasets that preserve document structure has limited progress. In this paper, we present our pipeline Wasm for processing the Common Crawl dataset to create a new Arabic multimodal dataset that uniquely provides markdown output. Unlike existing Arabic corpora that focus solely on text extraction, our approach preserves the structural integrity of web content while maintaining flexibility for both text-only and multimodal pre-training scenarios. We provide a comprehensive comparative analysis of our data processing pipeline against those used for major existing datasets, highlighting the convergences in filtering strategies and justifying our specific design choices. To support future research, we publicly release a representative dataset dump along with the multimodal processing pipeline for Arabic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。