arXiv:2608.23645cs.CLcs.LG2026-08

用嵌入向量验证乌尔都语动词的主-轻动词差异与关联

Contextual Embedding Evidence for Main--Light Verb Distinctions in Urdu

  • 通过三种BERT模型分析1126个句子,发现主用与轻用表示显著分离
  • 轻动词在掩码预测中仍可准确识别,最高准确率达0.866
  • 结果支持语法理论中主轻动词既有区别又保留词汇联系的观点

乌尔都语轻动词在表达事件结构时具有象征性意义,同时与对应主要动词保持词汇关联。本研究利用UrduBERT、DunbaaBERT和多语言BERT的上下文嵌入,在包含7个乌尔都语动词的1126个自然句中检验了Butt理论的表征预测。所有21个动词-模型组合中,主用与轻用均表现出显著表征分离;同时,同词干的主用与轻用中心点距离小于不同词干的主-轻配对,支持持续的词汇关联性。在仅限轻用的七分类预测任务中,目标词被掩码后,UrduBERT仍达到0.866准确率和0.852宏平均F1。在前形不重叠评估下,UrduBERT保持0.782准确率,表明其能泛化至未见的局部动词组合。这些发现为Butt的理论提供了计算证据:乌尔都语轻动词在系统上区别于主用,但仍保留词干和动词特异的表征结构。

原文摘要 · Abstract (English)

Urdu light verbs contribute schematic event-structural meaning while remaining lexically related to corresponding main verbs. This study tests representational predictions derived from Butt's analysis using contextual embeddings from UrduBERT, DunbaaBERT, and multilingual BERT across 1,126 naturally occurring sentences containing seven Urdu verbs. Main and light uses show significant representational separation in all 21 verb--model comparisons. At the same time, same-lemma main and light centroids are consistently closer than mismatched main--light lemma pairs, supporting continued lexical relatedness. In a seven-way prediction task restricted to light uses, verb identity remains recoverable after the target is masked, with UrduBERT achieving 0.866 accuracy and 0.852 macro-F1. UrduBERT also retains 0.782 accuracy under a preceding-form-disjoint evaluation, indicating generalization beyond repeated local verb combinations. These findings provide computational evidence consistent with Butt's account that Urdu light verbs differ systematically from their main uses while retaining lemma-specific and verb-specific representational structure.

语言学嵌入乌尔都语语法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。