首个支持八种语言与多种文字的像素级多语言生成模型
MIXAR: Scaling Autoregressive Pixel-based Language Models to Multiple Languages and Scripts
- 基于像素构建多语言模型,避免分词困扰
- 0.5B参数下在多语言任务中显著提升生成与判别性能
- 对未训练语言和文字攻击均具强鲁棒性,适合跨语言应用
像素基础语言模型正成为传统分词方法的替代方案,有望克服分词带来的挑战。然而,不同语言间固有的视觉差异给像素空间中的多语言泛化带来了重大障碍。本文提出MIXAR,首个在八种不同语言及多种书写系统上训练的生成式像素基础语言模型。我们通过实证评估MIXAR在判别与生成任务上优于以往像素模型及对比的分词模型。此外,模型对训练中未见的语言仍保持稳健。当模型规模扩展至0.5B参数时,不仅在生成任务(如LAMBADA)中表现更优,且在面对拼写攻击等输入扰动时也展现出更强鲁棒性。
原文摘要 · Abstract (English)
Pixel-based language models are gaining momentum as alternatives to traditional token-based approaches, promising to circumvent tokenization challenges. However, the inherent perceptual diversity across languages poses a significant hurdle for multilingual generalization in pixel space. This paper introduces MIXAR, the first generative pixel-based language model trained on eight different languages utilizing a range of different scripts. We empirically evaluate MIXAR against previous pixel-based models as well as comparable tokenizer-based models, demonstrating substantial performance improvement on discriminative and generative multilingual tasks. Additionally, we show how MIXAR is robust to languages never seen during the training. These results are further strengthened when scaling the model to 0.5B parameters which not only improves its capabilities in generative tasks like LAMBADA but also its robustness when challenged with input perturbations such as orthographic attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。