通过控制词汇频率,揭示大模型在罕见词下的鲁棒性与脆弱性
FreqBLiMP: Frequency-Controlled Minimal Pairs Reveal Robustness and Fragility of LLMs Under Lexical Rarity

- 构建频率可控的最小对测试集FreqBLiMP,保持语法对比不变
- 词汇越罕见,句子似然度越低,但整体判断准确率下降有限
- 模型对词形变化稳健,但依赖词项特异性信息的任务明显退化
最小对基准如BLiMP通过测试语言模型是否偏好可接受句而非仅语法上微小不同的不可接受句来评估语言知识。然而,这些基准大多忽略词汇频率的变化,而词汇频率是自然语言使用中普遍且高度偏斜的特性。因此,现有评估未能检验当对比涉及罕见词汇时,语法偏好是否仍稳定。我们提出FreqBLiMP,即对BLiMP的频率控制扩展版本,在明确的Zipf频率分布下重新生成全部67个句型范式,同时保持每个最小对的语法差异。在多个规模不同、开源权重的大型语言模型上进行评估发现:降低词汇频率导致句子似然度一致且单调下降,但总体对比接受度准确率仅略有下降。然而,这一整体稳定性掩盖了不同语言现象间的显著差异:模型在显式形态句法泛化任务上保持稳健,但在需要词项特异性信息的现象上显著退化。
原文摘要 · Abstract (English)
Minimal-pair benchmarks such as BLiMP evaluate linguistic knowledge by testing whether language models (LMs) prefer acceptable sentences over minimally different unacceptable ones. However, these benchmarks largely ignore lexical frequency variation, despite lexical frequency being a pervasive and highly skewed property of natural language use. Consequently, existing evaluations do not test whether grammatical preferences remain stable when contrasts involve rare lexical items. We introduce FreqBLiMP, a frequency-controlled extension of BLiMP that regenerates all 67 paradigms under explicit Zipf-frequency regimes while preserving each minimal-pair's grammatical contrast. Evaluating multiple open-weight LLM families across scales, we find that decreasing lexical frequency produces a consistent, monotonic decrease in sentence likelihood, but only a modest reduction in overall contrastive acceptability accuracy. However, this aggregate stability masks substantial variation across linguistic phenomena, with LLMs remaining robust on overt morphosyntactic generalization while degrading on phenomena that require lemma-specific information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。