构建超百万级多语言图像翻译数据集,提升模型真实场景适应能力
MIT-10M: A Large Scale Parallel Corpus of Multilingual Image Translation
- 基于真实数据构建1000万+图文对,覆盖14种语言与多难度任务
- 模型在该数据集上微调后性能提升3倍,显著优于基线
- 适合研究多语言视觉理解、跨模态翻译的开发者使用
图像翻译(IT)在多个领域具有巨大潜力,可将图像中的文本内容翻译为多种语言。然而,现有数据集在规模、多样性和质量上存在局限,制约了IT模型的发展与评估。为此,我们提出MIT-10M,一个大规模并行多语言图像翻译语料库,包含超过1000万张真实世界图像-文本对,经过大量数据清洗与多语言翻译验证。数据集涵盖84万张不同尺寸的图像,28个类别,三种难度级别的任务,以及14种语言的图文对,相较已有数据集有显著提升。我们在该数据集上开展广泛实验,结果表明其在评估模型应对复杂现实图像翻译任务时具备更强适应性。此外,经MIT-10M微调的模型性能较基线提升三倍,进一步验证其优越性。
原文摘要 · Abstract (English)
Image Translation (IT) holds immense potential across diverse domains, enabling the translation of textual content within images into various languages. However, existing datasets often suffer from limitations in scale, diversity, and quality, hindering the development and evaluation of IT models. To address this issue, we introduce MIT-10M, a large-scale parallel corpus of multilingual image translation with over 10M image-text pairs derived from real-world data, which has undergone extensive data cleaning and multilingual translation validation. It contains 840K images in three sizes, 28 categories, tasks with three levels of difficulty and 14 languages image-text pairs, which is a considerable improvement on existing datasets. We conduct extensive experiments to evaluate and train models on MIT-10M. The experimental results clearly indicate that our dataset has higher adaptability when it comes to evaluating the performance of the models in tackling challenging and complex image translation tasks in the real world. Moreover, the performance of the model fine-tuned with MIT-10M has tripled compared to the baseline model, further confirming its superiority.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。