探究位置编码对不同语言结构的影响,发现无明显关联
On the Interplay between Positional Encodings, Morphological Complexity, and Word Order Flexibility
- 对比绝对、相对和无位置编码的模型在七种语言上的表现
- 未发现形态复杂度与语序灵活性之间存在预期交互作用
- 结果提示任务类型和评估指标对结论稳定性至关重要
语言模型架构多以英语为基础构建,随后应用于其他语言。这种架构偏向是否导致结构差异大的语言性能下降仍存疑问。本文聚焦位置编码这一设计选择,基于形态复杂度与语序灵活性之间的权衡假说进行研究。该假说认为二者呈此消彼长关系:形态越复杂,语序越灵活。我们为七种类型学多样化的语言预训练了采用绝对、相对及无位置编码的单语模型,并在四个下游任务上进行评估。结果表明,与先前研究相反,未观察到位置编码与形态复杂度或语序灵活性之间的显著交互作用,其影响依赖于具体任务、语言和评估指标,强调了结论稳定性的关键因素。
原文摘要 · Abstract (English)
Language model architectures are predominantly first created for English and subsequently applied to other languages. It is an open question whether this architectural bias leads to degraded performance for languages that are structurally different from English. We examine one specific architectural choice: positional encodings, through the lens of the trade-off hypothesis: the supposed interplay between morphological complexity and word order flexibility. This hypothesis posits a trade-off between the two: a more morphologically complex language can have a more flexible word order, and vice-versa. Positional encodings are a direct target to investigate the implications of this hypothesis in relation to language modelling. We pretrain monolingual model variants with absolute, relative, and no positional encodings for seven typologically diverse languages and evaluate them on four downstream tasks. Contrary to previous findings, we do not observe a clear interaction between position encodings and morphological complexity or word order flexibility, as measured by various proxies. Our results show that the choice of tasks, languages, and metrics are essential for drawing stable conclusions
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。