证明多语言嵌入空间不存在理论上的性能衰减问题
There is No Theoretical Curse of Multilinguality For Embedding Space Structure
- 从理论上定义完美多语言性的两个条件
- 证明所需维度仅随语言数对数增长
- 揭示实际性能下降源于数据与训练条件
多语言自然语言处理的核心目标是实现每种语言的高单语性能,同时在大规模语言覆盖下保持跨语言对齐。多语言诅咒指随着语言数量增加,多语言模型性能下降的现象,威胁这一目标。本文探讨嵌入空间是否天然无法在不显著增加容量的前提下实现完美多语言性。我们首先形式化了‘完美多语言性’的两个条件,进而证明实现该目标所需的最小维度仅随语言数量对数增长。这表明嵌入空间结构上并不存在多语言诅咒。该结果暗示实证中观察到的诅咒源于真实世界的数据与训练条件。我们通过小规模实证研究验证了这一理解。本论文首次从理论和内在角度剖析多语言诅咒,为理解该现象提供科学依据。
原文摘要 · Abstract (English)
A central goal of multilingual NLP is to achieve high monolingual performance per language and cross-lingual alignment for large-scale language coverage with a multilingual model. The curse of multilinguality describes the phenomenon of degradation in multilingual model performance as we increase language coverage, posing a threat to the above goal. This paper asks whether multilingual embedding spaces are inherently incapable of achieving perfect multilinguality without a prohibitive increase in required capacity. We first formalize the goal of "perfect multilinguality", embodied in two multilinguality conditions. We then prove that the minimum dimensionality required for perfect multilinguality grows only logarithmically in the number of languages. That is, we show that there is no theoretical curse of multilinguality for embedding space structure. This suggests that the empirical curse of multilinguality is a result of real world data and training conditions. We back this understanding with a small-scale empirical study. Our paper provides the first theoretical and intrinsic perspective on the curse of multilinguality, with implications for the scientific understanding of this phenomenon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。