arXiv:2601.20592cs.CL2026-01被引 1

研究波斯语的接触痕迹,发现模型中形态特征受语言影响更大。

A Computational Approach to Language Contact -- A Case Study of Persian

  • 用单语模型中间层表征探测语言接触影响
  • 形态特征如格和性别明显受接触影响
  • 适合关注语言演化与模型表征的研究者

我们研究单语语言模型中间表示中语言接触的结构痕迹。以历史上接触频繁的波斯语(法尔斯语)为例,探测在不同接触程度和类型的语言刺激下,经过波斯语训练的模型中间表示中的信息编码情况。方法量化了中间表示中编码的语言信息量,并评估其在不同句法特征上的分布。结果表明,普遍句法信息对历史接触不敏感,而形态特征如格和性别则显著受语言特异性结构影响,说明单语模型中的接触效应具有选择性且结构受限。

原文摘要 · Abstract (English)

We investigate structural traces of language contact in the intermediate representations of a monolingual language model. Focusing on Persian (Farsi) as a historically contact-rich language, we probe the representations of a Persian-trained model when exposed to languages with varying degrees and types of contact with Persian. Our methodology quantifies the amount of linguistic information encoded in intermediate representations and assesses how this information is distributed across model components for different morphosyntactic features. The results show that universal syntactic information is largely insensitive to historical contact, whereas morphological features such as Case and Gender are strongly shaped by language-specific structure, suggesting that contact effects in monolingual language models are selective and structurally constrained.

语言模型形态学语言接触

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。