构建11种印度语多模态数据集,助力本土化视觉语言模型训练
Chitrakshara: A Large Multilingual Multimodal Dataset for Indian languages
- 从Common Crawl采集11种印度语多模态数据,构建双子集
- 包含19300万图像、300亿文本词元与4400万图文对
- 适配文化多样性需求,推动非英语VLM研究发展
多模态研究长期聚焦单图理解,对多图场景探索不足。现有模型虽通过大规模交错图文数据预训练提升多图理解能力,但多数视觉语言模型(VLM)仍以英文数据为主,难以覆盖印度语种。为此,我们推出Chitrakshara数据集系列,涵盖11种印度语言,数据源自Common Crawl。该系列包含:(1) Chitrakshara-IL,大规模交错预训练数据集,含1.93亿图像、300亿文本词元及5000万多语言文档;(2) Chitrakshara-Cap,包含4400万图文对与7.33亿词元。本文详述数据采集流程,包括筛选、清洗与处理方法,并开展全面的质量与多样性分析,评估其在印地语系语言中的代表性,验证其在开发更具文化包容性的VLM方面的潜力。
原文摘要 · Abstract (English)
Multimodal research has predominantly focused on single-image reasoning, with limited exploration of multi-image scenarios. Recent models have sought to enhance multi-image understanding through large-scale pretraining on interleaved image-text datasets. However, most Vision-Language Models (VLMs) are trained primarily on English datasets, leading to inadequate representation of Indian languages. To address this gap, we introduce the Chitrakshara dataset series, covering 11 Indian languages sourced from Common Crawl. It comprises (1) Chitrakshara-IL, a large-scale interleaved pretraining dataset with 193M images, 30B text tokens, and 50M multilingual documents, and (2) Chitrakshara-Cap, which includes 44M image-text pairs with 733M tokens. This paper details the data collection pipeline, including curation, filtering, and processing methodologies. Additionally, we present a comprehensive quality and diversity analysis to assess the dataset's representativeness across Indic languages and its potential for developing more culturally inclusive VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。