研究像素语言模型在低资源语言中的视觉适应,发现书写系统相似性提升迁移效果。
Seeing the Unseen: Visual Similarity for Pixel Language Model Adaptation

- 引入四种渲染级视觉相似性度量,评估不同文字系统的视觉接近程度
- 在藏文上验证:书写系统越相似,语义迁移越强,即使数据极少也有效
- 单语模型经混合脚本微调,在句子级任务上表现优于多语预训练模型
基于像素的语言模型(LMs)通过处理文本的图像来替代传统分词器,其跨语言迁移高度依赖文字系统的视觉与结构特性。然而,针对具有复杂形态、使用独特脚本的低资源语言,这类模型的适应机制尚未被深入探索。以藏文为例,我们分析了持续预训练受数据规模、初始脚本暴露程度及来自其他婆罗米系文字的跨语言迁移的影响。引入四种渲染级度量来量化视觉脚本相似性,并在三个下游任务上进行评估。结果表明,更高的正字法相近性可增强语义迁移,即便在严重数据约束下依然有效。此外,性能存在起点依赖性:多语言预训练模型PIXEL-M4虽初始表现更好,但后续适应能力受限;而从单语模型PIXEL出发,结合多种脚本微调,反而在句子级任务上获得更大提升。这些度量与案例研究为类似低资源场景下的数据选择和脚本适配提供了实证依据。
原文摘要 · Abstract (English)
Pixel-based language models (LMs) replace traditional tokenizers by processing rendered images of text, making cross-lingual transfer heavily dependent on the visual and structural properties of writing systems. However, the dynamics of adapting these models to low-resource languages with complex morphology and written in unique scripts are not yet explored. Using Tibetan as a case study, we analyze how continued pre-training of pixel-based LMs is influenced by data scale, initial script exposure, and cross-lingual transfer from languages written in other Brahmic scripts. We introduce four rendering-level metrics to quantify visual script similarity. We evaluate downstream performance across three tasks. Our results show that higher orthographic proximity enhances semantic transfer, even under severe data constraints. Additionally, we find a performance asymmetry based on the pre-training starting point: while multilingual pre-training PIXEL-M4 has stronger initial performance, its capacity for subsequent adaptation seems to be constrained, whereas adapting a monolingual model PIXEL with mixed scripts yields more gains on sentence-level tasks. Our metrics and case study offer empirical observations that could help inform data selection and script adaptation choices when working with pixel-based models in similar low-resource settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。