检验词嵌入对齐度指标BLI的可靠性,发现其在某些场景下失效并提出改进方法。
How Good is BLI as an Alignment Measure: A Study in Word Embedding Paradigm
- 提出基于词干的BLI方法,更好捕捉语言屈折特征
- 发现组合对齐技术多数情况更优,低资源语言中多语言模型反而占优
- 设计更有效的词汇修剪策略,提升对齐评估精度
随着低资源领域单语嵌入研究减少,多语言嵌入因其能处理混合语言、实现无语言依赖文档处理,且避免单语嵌入对齐难题,已成为主流选择。但这一优势是否全面?多语言模型是否在所有方面均优于对齐后的单语模型?其更高计算成本是否总被证明合理?是否存在折中方案?双语词典归纳(BLI)是衡量两嵌入空间对齐程度最广泛使用的指标。本研究探讨了BLI作为对齐度量的优缺点。进一步评估传统对齐技术、新型多语言模型及组合对齐技术在高资源与低资源语言场景下的BLI表现。同时考察语言家族对的影响。结果表明,某些情况下BLI未能真实反映对齐程度,并提出相应解决方案。本文提出一种基于词干的新型BLI方法,考虑语言的屈折特性,优于主流词级方法。此外引入一种更具信息量的词汇修剪技术,尤其适用于多语言嵌入模型的BLI评估。实验显示,组合对齐技术总体表现更佳,但在低资源语言场景下,多语言嵌入模型表现更优。
原文摘要 · Abstract (English)
Sans a dwindling number of monolingual embedding studies originating predominantly from the low-resource domains, it is evident that multilingual embedding has become the de facto choice due to its adaptability to the usage of code-mixed languages, granting the ability to process multilingual documents in a language-agnostic manner, as well as removing the difficult task of aligning monolingual embeddings. But is this victory complete? Are the multilingual models better than aligned monolingual models in every aspect? Can the higher computational cost of multilingual models always be justified? Or is there a compromise between the two extremes? Bilingual Lexicon Induction is one of the most widely used metrics in terms of evaluating the degree of alignment between two embedding spaces. In this study, we explore the strengths and limitations of BLI as a measure to evaluate the degree of alignment of two embedding spaces. Further, we evaluate how well traditional embedding alignment techniques, novel multilingual models, and combined alignment techniques perform BLI tasks in the contexts of both high-resource and low-resource languages. In addition to that, we investigate the impact of the language families to which the pairs of languages belong. We identify that BLI does not measure the true degree of alignment in some cases and we propose solutions for them. We propose a novel stem-based BLI approach to evaluate two aligned embedding spaces that take into account the inflected nature of languages as opposed to the prevalent word-based BLI techniques. Further, we introduce a vocabulary pruning technique that is more informative in showing the degree of the alignment, especially performing BLI on multilingual embedding models. Often, combined embedding alignment techniques perform better while in certain cases multilingual embeddings perform better (mainly low-resource language cases).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。