arXiv:2410.12011cs.CL2024-10EMNLP被引 2

探究像素级语言模型的视觉与语言能力边界

Pixology: Probing the Linguistic and Visual Capabilities of Pixel-based Language Models

  • 通过多任务探针分析模型各层的视觉与语言表征
  • 发现模型高层才逐渐具备语义抽象能力,底层仅学表面视觉特征
  • 输入端加入书写规范可加速表层特征学习,适合模型设计者参考

像素级语言模型作为子词模型的替代方案,尤其在处理多种文字系统方面表现突出。以PIXEL为例,这是一种基于视觉变换器、在渲染文本上预训练的模型,虽在跨文字迁移和抗字形扰动方面表现良好,但在多数语言任务中仍不如单语子词模型(如BERT)。这一差距引发对其语言知识习得程度的质疑:其性能更多来自视觉能力还是语言理解?为此,我们通过多种语言与视觉任务对PIXEL进行探针分析,评估其在视觉到语言谱系中的位置。结果表明,模型在视觉与语言理解之间存在显著差距:低层主要捕捉浅层视觉特征,高层才逐步学习句法与语义抽象。此外,我们还考察了不同文本渲染策略训练的PIXEL变体,发现输入端引入特定字形约束能促进早期表面特征的学习。本研究旨在为像素级语言模型的进一步发展提供洞见。

原文摘要 · Abstract (English)

Pixel-based language models have emerged as a compelling alternative to subword-based language modelling, particularly because they can represent virtually any script. PIXEL, a canonical example of such a model, is a vision transformer that has been pre-trained on rendered text. While PIXEL has shown promising cross-script transfer abilities and robustness to orthographic perturbations, it falls short of outperforming monolingual subword counterparts like BERT in most other contexts. This discrepancy raises questions about the amount of linguistic knowledge learnt by these models and whether their performance in language tasks stems more from their visual capabilities than their linguistic ones. To explore this, we probe PIXEL using a variety of linguistic and visual tasks to assess its position on the vision-to-language spectrum. Our findings reveal a substantial gap between the model's visual and linguistic understanding. The lower layers of PIXEL predominantly capture superficial visual features, whereas the higher layers gradually learn more syntactic and semantic abstractions. Additionally, we examine variants of PIXEL trained with different text rendering strategies, discovering that introducing certain orthographic constraints at the input level can facilitate earlier learning of surface-level features. With this study, we hope to provide insights that aid the further development of pixel-based language models.

视觉语言模型语言探针像素建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。