数据决定语言模型的语序偏好,SVO语言因资源多而占优。
Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models
- 用人工语言和真实语言对比模型语序偏好,发现数据驱动倾向。
- 小规模时无明显偏好,大数据下更倾向SVO结构,即便SOV更普遍。
- 适合关注大模型偏见、语言多样性风险的研究者与从业者。
我们系统比较了192种人工语言和类型多样化的自然语言中解码器仅有的语言模型的语序偏好。在人工语言中,模型表现出左嵌套偏好,这既不符合自然语言普遍性,也不符合人类学习词序的倾向。在自然语言中,单语模型在小规模时无明显基础语序偏好;但随着数据量增加,对右嵌套主谓宾(SVO)语言的偏好逐渐显现,尽管跨语言中主宾谓(SOV)是最常见的结构。这一SVO优势也延伸至多语模型,且与语言资源水平和数据质量相关,而非语序本身。因此,同一架构在人工与自然语言上表现出相反偏好,表明实际观察到的语序偏好是数据驱动的。由于高资源语言几乎都是SVO,这种偏好可能逐步削弱语序多样性,尤其对那些能灵活使用多种语序的语言构成风险,随着大模型的广泛应用。
原文摘要 · Abstract (English)
We systematically compare word order preferences in decoder-only language models across 192 artificial languages and typologically diverse natural languages. On artificial languages, models exhibit a left-branching preference that aligns with neither natural language universals nor human word order learning biases. On natural languages, monolingual models show no clear base word order bias at small scales, but as data grows, a preference for right-branching subject-verb-object (SVO) languages emerges while SOV falls behind despite being the most frequent order cross-linguistically. This SVO advantage extends to multilingual models and correlates with language resource level and data quality rather than word order. Thus, the same architecture exhibits opposite preferences on artificial and natural languages, establishing that word order biases observed in practice are data-driven. Since highly-resourced languages are overwhelmingly SVO, these biases risk gradually reducing word order diversity, particularly in languages that productively use multiple word orders, with the widespread adoption of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。