评测大模型对格鲁吉亚语格标记的精准度,发现对格标记表现最差。
Targeted Syntactic Evaluation of Language Models on Georgian Case Alignment
- 用树库生成最小对比句,构建7个任务共370个语法测试
- 模型在三种格中准确率最低的是格标记(<50%),最高是主格
- 数据稀缺与格标记特殊性共同导致模型表现不佳
本文评估了基于Transformer的语言模型在格鲁吉亚语分裂作通格标记系统中的表现,该系统极为罕见,通过主格、作格和与格三种名词形式区分主语和宾语。采用基于树库的Grew查询语言生成最小对比句,构建包含7个任务、每任务50-70个样本的370条语法测试集,每个样本测试三种名词形式。评估五种编码器及两种仅解码器模型,使用词级和/或句级准确率。结果显示,无论具体句法结构如何,模型在作格标记上的表现最差,主格标记最优,准确率排序为:主格 > 与格 > 作格。性能与三类形式的整体频次分布一致(NOM > DAT > ERG)。尽管低资源语言数据稀缺是已知问题,但作格角色的高度特异性与训练数据不足可能是其表现不佳的关键原因。数据集已公开,方法可推广至其他缺乏基准的语言。
原文摘要 · Abstract (English)
This paper evaluates the performance of transformer-based language models on split-ergative case alignment in Georgian, a particularly rare system for assigning grammatical cases to mark argument roles. We focus on subject and object marking determined through various permutations of nominative, ergative, and dative noun forms. A treebank-based approach for the generation of minimal pairs using the Grew query language is implemented. We create a dataset of 370 syntactic tests made up of seven tasks containing 50-70 samples each, where three noun forms are tested in any given sample. Five encoder- and two decoder-only models are evaluated with word- and/or sentence-level accuracy metrics. Regardless of the specific syntactic makeup, models performed worst in assigning the ergative case correctly and strongest in assigning the nominative case correctly. Performance correlated with the overall frequency distribution of the three forms (NOM > DAT > ERG). Though data scarcity is a known issue for low-resource languages, we show that the highly specific role of the ergative along with a lack of available training data likely contributes to poor performance on this case. The dataset is made publicly available and the methodology provides an interesting avenue for future syntactic evaluations of languages where benchmarks are limited.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。