arXiv:2508.16867cs.CL2025-08EMNLP被引 3

构建法语魁北克方言语法可接受性数据集,评估大模型语言判断能力

QFrCoLA: a Quebec-French Corpus of Linguistic Acceptability Judgments

  • 基于25,153条领域内/2,675条领域外句子构建法语魁北克方言可接受性数据集
  • 微调后的Transformer模型在该数据集上表现最优,零样本大模型表现差
  • 数据集反映语言规范而非主观感受,适合评估模型语言知识掌握程度

大型基于Transformer的語言模型在下游任务中表現出色,但其內部如何掌握語言知識仍不清楚。為此,本文提出QFrCoLA(魁北克法語語言可接受性判斷語料庫),一個包含25,153條領域內和2,675條領域外句子的二元可接受性判斷數據集。研究利用該數據集與另外七個語言學二元可接受性判斷語料庫,對七種語言模型進行基線測試。結果顯示,微調後的Transformer模型在多數語言上表現良好;而零樣本大語言模型在此任務上表現不佳。針對QFrCoLA,微調的Transformer模型平均表現優於其他方法。此外,預訓練的跨語言大模型未在預訓練階段習得魁北克法語的語言判斷能力。實驗表明,本數據集基於語言規範而非個人主觀感受,具有挑戰性,能有效評估模型的語言判斷能力。

原文摘要 · Abstract (English)

Large and Transformer-based language models perform outstandingly in various downstream tasks. However, there is limited understanding regarding how these models internalize linguistic knowledge, so various linguistic benchmarks have recently been proposed to facilitate syntactic evaluation of language models across languages. This paper introduces QFrCoLA (Quebec-French Corpus of Linguistic Acceptability Judgments), a normative binary acceptability judgments dataset comprising 25,153 in-domain and 2,675 out-of-domain sentences. Our study leverages the QFrCoLA dataset and seven other linguistic binary acceptability judgment corpora to benchmark seven language models. The results demonstrate that, on average, fine-tuned Transformer-based LM are strong baselines for most languages and that zero-shot binary classification large language models perform poorly on the task. However, for the QFrCoLA benchmark, on average, a fine-tuned Transformer-based LM outperformed other methods tested. It also shows that pre-trained cross-lingual LLMs selected for our experimentation do not seem to have acquired linguistic judgment capabilities during their pre-training for Quebec French. Finally, our experiment results on QFrCoLA show that our dataset, built from examples that illustrate linguistic norms rather than speakers' feelings, is similar to linguistic acceptability judgment; it is a challenging dataset that can benchmark LM on their linguistic judgment capabilities.

语言模型语法评测法语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。