多语言模型在词义消歧上表现差,因容量受限导致语义表征弱化。
Capacity Constraints and the Multilingual Penalty for Lexical Disambiguation
- 用人工标注数据对比单语与多语模型的词义消歧能力
- 多语模型在三项容量限制上均表现更差,性能下降可被这些因素解释
- 适合关注多语言模型瓶颈的研究者阅读
多语言语言模型有时表现不如单语模型,可能源于容量限制。我们通过控制数据集量化了词义消歧任务中的‘多语言惩罚’——该任务需要精确的语义表示和上下文建模。使用英语和西班牙语中歧义词的人类相关性判断数据集,对比同一系列的单语与多语模型,发现多语模型表现持续偏低。进一步分析三种潜在容量约束:表征(嵌入正交性下降)、注意力(对消歧线索关注度降低)、词汇(更多多标记分段)。多语模型在三方面均有证据显示受限,且这些因素能统计解释原归因于多语身份的性能差异。结果表明多语模型确实面临多重容量约束,且与消歧性能下降相关。
原文摘要 · Abstract (English)
Multilingual language models (LMs) sometimes under-perform their monolingual counterparts, possibly due to capacity limitations. We quantify this ``multilingual penalty'' for lexical disambiguation--a task requiring precise semantic representations and contextualization mechanisms--using controlled datasets of human relatedness judgments for ambiguous words in both English and Spanish. Comparing monolingual and multilingual LMs from the same families, we find consistently reduced performance in multilingual LMs. We then explore three potential capacity constraints: representational (reduced embedding isotropy), attentional (reduced attention to disambiguating cues), and vocabulary-related (increased multi-token segmentation). Multilingual LMs show some evidence of all three limitations; moreover, these factors statistically account for the variance formerly attributed to a model's multilingual status. These findings suggest both that multilingual LMs do suffer from multiple capacity constraints, and that these constraints correlate with reduced disambiguation performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。