arXiv:2505.21265cs.CLcs.AI2025-05EMNLP被引 15

多语言预训练让图像语言模型更好支持非拉丁文字

Multilingual Pretraining for Pixel Language Models

  • 在英、印、乌、中四种语言图像上联合预训练
  • 对非拉丁文字任务表现优于单语模型,跨语言语义空间更对齐
  • 适合需要多语言文本理解的视觉语言研究者

像素语言模型直接处理渲染文本图像,无需固定词表。尽管该类模型在跨语言迁移中表现优异,但多语言预训练仍较少被探索。本文提出 PIXEL-M4,基于英文、印地文、乌克兰文和简体中文四种语言的图像进行预训练。在语义与句法任务上的多语言评估显示,PIXEL-M4 在非拉丁文字上的表现优于仅使用英文预训练的模型。词级探针分析表明,即使在预训练中未出现的语言中,模型也能捕捉丰富的语言特征。隐藏表示分析进一步揭示,多语言预训练使不同语言的语义嵌入空间高度对齐。该工作证明,多语言预训练能显著提升像素语言模型对多样语言的支持能力。

原文摘要 · Abstract (English)

Pixel language models operate directly on images of rendered text, eliminating the need for a fixed vocabulary. While these models have demonstrated strong capabilities for downstream cross-lingual transfer, multilingual pretraining remains underexplored. We introduce PIXEL-M4, a model pretrained on four visually and linguistically diverse languages: English, Hindi, Ukrainian, and Simplified Chinese. Multilingual evaluations on semantic and syntactic tasks show that PIXEL-M4 outperforms an English-only counterpart on non-Latin scripts. Word-level probing analyses confirm that PIXEL-M4 captures rich linguistic features, even in languages not seen during pretraining. Furthermore, an analysis of its hidden representations shows that multilingual pretraining yields a semantic embedding space closely aligned across the languages used for pretraining. This work demonstrates that multilingual pretraining substantially enhances the capability of pixel language models to effectively support a diverse set of languages.

像素语言模型多语言跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。