研究土耳其语轻动词构式分类的关键信号,发现形态语法不足、词汇身份敏感。
From Lemmas to Dependencies: What Signals Drive Light Verbs Classification?
- 通过限制输入,对比词元、语法和完整特征的表现
- 仅靠粗粒度形态语法无法稳健识别轻动词构式
- 词元信息虽有用但依赖归一化处理,需谨慎设计
轻动词构式(LVCs)是极具挑战性的动词多词表达,尤其在土耳其语中,丰富的形态变化和产物性复合谓词导致习语义与字面动词-论元用法之间的对比极小。本文系统性地限制模型输入,探究驱动LVC分类的关键信号。基于UD标注的监督,比较了词元驱动基线(词元TF-IDF + 逻辑回归;BERTurk训练于词元序列)、仅语法的逻辑回归(基于UD形态句法:UPOS/DEPREL/MORPH)以及全输入BERTurk基线。在包含随机负例、词汇对照(NLVC)和LVC正例的受控诊断集上评估,报告分项性能以揭示决策边界行为。结果表明,在受控对比下,仅靠粗粒度形态语法不足以实现鲁棒的LVC检测;而词汇身份虽支持判断,但对校准与归一化选择敏感。总体而言,研究推动针对土耳其语多词表达的针对性评估,并指出‘仅词元’并非单一明确表征,其效果高度依赖归一化的具体实现方式。
原文摘要 · Abstract (English)
Light verb constructions (LVCs) are a challenging class of verbal multiword expressions, especially in Turkish, where rich morphology and productive complex predicates create minimal contrasts between idiomatic predicate meanings and literal verb--argument uses. This paper asks what signals drive LVC classification by systematically restricting model inputs. Using UD-derived supervision, we compare lemma-driven baselines (lemma TF--IDF + Logistic Regression; BERTurk trained on lemma sequences), a grammar-only Logistic Regression over UD morphosyntax (UPOS/DEPREL/MORPH), and a full-input BERTurk baseline. We evaluate on a controlled diagnostic set with Random negatives, lexical controls (NLVC), and LVC positives, reporting split-wise performance to expose decision-boundary behavior. Results show that coarse morphosyntax alone is insufficient for robust LVC detection under controlled contrasts, while lexical identity supports LVC judgments but is sensitive to calibration and normalization choices. Overall, Our findings motivate targeted evaluation of Turkish MWEs and show that ``lemma-only'' is not a single, well-defined representation, but one that depends critically on how normalization is operationalized.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。