arXiv:2506.01846cs.CL2025-06ACL被引 7

语法足以解释双语者换语位置偏好,且模型可泛化到未见语言对。

Code-Switching and Syntax: A Large-Scale Experiment

  • 仅用语法信息预测换语位置,无需语义或词汇干扰。
  • 模型区分换语句的准确率与母语者相当,达85%以上。
  • 学习的语法模式能跨语言推广至未见过的语言组合。

理论上的换语研究虽揭示了双语者在句中特定位置更频繁换语的现象,但缺乏大规模、多语言、跨现象的实验验证。本文设计一项实验,确保预测系统仅依赖句法信息。结果表明,仅凭语法特征,自动系统即可达到与双语人类相当的水平,准确区分最小对例中的换语句子(准确率超过85%)。此外,所学得的句法模式具有良好的跨语言泛化能力,适用于未见过的语言配对。该发现支持换语主要由语言句法结构决定的观点。

原文摘要 · Abstract (English)

The theoretical code-switching (CS) literature provides numerous pointwise investigations that aim to explain patterns in CS, i.e. why bilinguals switch language in certain positions in a sentence more often than in others. A resulting consensus is that CS can be explained by the syntax of the contributing languages. There is however no large-scale, multi-language, cross-phenomena experiment that tests this claim. When designing such an experiment, we need to make sure that the system that is predicting where bilinguals tend to switch has access only to syntactic information. We provide such an experiment here. Results show that syntax alone is sufficient for an automatic system to distinguish between sentences in minimal pairs of CS, to the same degree as bilingual humans. Furthermore, the learnt syntactic patterns generalise well to unseen language pairs.

换语研究句法分析双语处理语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。