arXiv:2411.02528cs.CL2024-11NAACL被引 13

提出新方法提升大模型与人类可接受性判断的匹配度

What Goes Into a LM Acceptability Judgment? Rethinking the Impact of Frequency and Length

  • 基于数据学习长度和词频调整参数,动态校正模型评分偏差
  • 大模型对词频敏感度更低,但需适度调整仍能更好匹配人类判断
  • 揭示大模型更善预测罕见词,是其抗频率干扰的关键原因

在用语言模型(LM)概率比较人类与模型的语言能力时,序列长度和词汇的单字频次对模型概率影响显著,而人类对此类因素相对鲁棒。以往研究对不同模型统一处理这些影响,假设各模型需相同程度的校正。本文提出MORCELA——一种从数据中学习长度与词频调整参数的新链接理论,使模型分数更贴近人类可接受性判断。实验表明,MORCELA在Pythia与OPT两类Transformer模型上均优于经典方法SLOR(Pauls and Klein, 2012; Lau et al., 2017)。进一步分析发现,SLOR对长度与词频的校正过度,且大模型对词频的依赖更低,但仍需一定调整。此外,大模型对罕见词的上下文预测能力更强,解释了其较低的词频敏感性。

原文摘要 · Abstract (English)

When comparing the linguistic capabilities of language models (LMs) with humans using LM probabilities, factors such as the length of the sequence and the unigram frequency of lexical items have a significant effect on LM probabilities in ways that humans are largely robust to. Prior works in comparing LM and human acceptability judgments treat these effects uniformly across models, making a strong assumption that models require the same degree of adjustment to control for length and unigram frequency effects. We propose MORCELA, a new linking theory between LM scores and acceptability judgments where the optimal level of adjustment for these effects is estimated from data via learned parameters for length and unigram frequency. We first show that MORCELA outperforms a commonly used linking theory for acceptability - SLOR (Pauls and Klein, 2012; Lau et al. 2017) - across two families of transformer LMs (Pythia and OPT). Furthermore, we demonstrate that the assumed degrees of adjustment in SLOR for length and unigram frequency overcorrect for these confounds, and that larger models require a lower relative degree of adjustment for unigram frequency, though a significant amount of adjustment is still necessary for all models. Finally, our subsequent analysis shows that larger LMs' lower susceptibility to frequency effects can be explained by an ability to better predict rarer words in context.

语言模型可接受性判断词频效应模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。