用语义相似的子词合并提升多语言模型理解能力
Tomato, Tomahto, Tomate: Do Multilingual Language Models Understand Based on Subword-Level Semantic Concepts?
- 将语义相近的子词及其嵌入合并为语义标记
- 在五项跨语言任务中,新方法零样本表现媲美甚至优于原模型
- 适用于不同分词器和模型规模的多语言模型研究
人类对文本的理解依赖于词语的通用语义概念,而非表面形式。当前多语言语言模型(mLMs)在多大程度上具备这种基于子词级语义概念的理解能力?本文通过合并语义相似的子词及其嵌入,构建“语义标记”,并在五个异构多语言下游任务上评估更新后的mLMs。结果表明,共享的通用语义可显著提升不同分词器和模型规模下的预测性能。对合并后子词的分析显示,其语义相似性涵盖多种语言和文字间的同义词与翻译关系。此外,在部分分类任务上,使用语义标记的零样本结果与原模型相当甚至更优,说明共享的子词级语义可能成为跨语言迁移的锚点。
原文摘要 · Abstract (English)
Human understanding of text depends on general semantic concepts of words rather than their superficial forms. To what extent does our human intuition transfer to language models? In this work, we study the degree to which current multilingual language models (mLMs) understand based on subword-level semantic concepts. To this end, we form "semantic tokens" by merging the semantically similar subwords and their embeddings, and evaluate the updated mLMs on five heterogeneous multilingual downstream tasks. Results show that the general shared semantics could get the models a long way in making the predictions on mLMs with different tokenizers and model sizes. Inspections of the grouped subwords show that they exhibit a wide range of semantic similarities, including synonyms and translations across many languages and scripts. Lastly, we find that the zero-shot results with semantic tokens are on par with or even better than the original models on certain classification tasks, suggesting that the shared subword-level semantics may serve as the anchors for cross-lingual transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。