arXiv:2411.06096cs.CL2024-11Transactions of th…被引 13

构建中文最小对基准,揭示大模型在指代等语法上的学习瓶颈。

A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese

  • 设计新指标SLLN-LP,解决句子长度差异带来的评估偏差。
  • 320亿参数模型仍难以处理中文指代、量词和省略现象。
  • 适合关注中文语言模型评测与语法能力的研究者。

本文提出ZhoBLiMP,首个覆盖100余种汉语语法现象的最小对基准,涵盖话题化、把字句等。我们从零训练不同分词器、参数量与语料量的中文语言模型,研究其学习曲线。为缓解最小对中句子长度不均导致的偏差,提出子线性长度归一化对数概率(SLLN-LP)指标。实验表明,即使32B参数模型在指代、量词和省略任务上仍有困难,且SLLN-LP有效缓解了ZhoBLiMP、JBLiMP和BLiMP中的长度偏差问题。结论强调未来评测需更精细设计,考虑链接函数、模型与最小对之间的复杂关系。

原文摘要 · Abstract (English)

We present ZhoBLiMP, the largest linguistic minimal pair benchmark for Chinese, with over 100 paradigms, ranging from topicalization to the \textit{Ba} construction. We then train from scratch a suite of Chinese language models (LMs) with different tokenizers, parameter sizes, and token volumes, to study the learning curves of LMs on Chinese. To mitigate the biases introduced by unequal lengths of the sentences in a minimal pair, we propose a new metric named sub-linear length normalized log-probabilities (SLLN-LP). Using SLLN-LP as the metric, our results show that \textsc{Anaphor}, \textsc{Quantifiers}, and \textsc{Ellipsis} in Chinese are difficult for LMs even up to 32B parameters, and that SLLN-LP successfully mitigates biases in ZhoBLiMP, JBLiMP and BLiMP. We conclude that future evaluations should be more carefully designed to consider the intricate relations between linking functions, LMs, and targeted minimal pairs.

语言模型中文评测语法能力最小对

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。