测试小模型能否理解对话中的隐含意义,发现数据多一点效果更好。
Are BabyLMs Deaf to Gricean Maxims? A Pragmatic Evaluation of Sample-efficient Language Models
- 用格里森会话准则设计新评测基准,检验小模型对言外之意的识别能力。
- 100M以下训练数据的模型比10M以下表现更好,但仍不及儿童和大模型。
- 适合关注小模型语用能力、高效语言模型研究的人阅读。
隐含意义是人类交流的核心,语言模型必须具备识别与解读能力。格里森(1975)提出一组会话准则,指导合作对话:说话人可能故意违反这些原则以传达字面之外的意义,听者则通过识别违规来推断言外之意。借鉴Surian等(1996)对儿童敏感性的研究,我们引入一个新基准,测试在少于10M和少于100M tokens上预训练的语言模型是否能区分遵守与违反格里森准则的语句。我们在五个准则上对比这些“婴儿模型”(BabyLMs),并与儿童及在3T tokens上预训练的大语言模型(LLM)进行比较。结果表明,100M以下训练的数据模型整体优于10M以下,但仍显著落后于儿童水平和LLM。说明适度增加数据可提升部分语用行为,使模型对语用维度的区分更精细。
原文摘要 · Abstract (English)
Implicit meanings are integral to human communication, making it essential for language models to be capable of identifying and interpreting them. Grice (1975) proposed a set of conversational maxims that guide cooperative dialogue, noting that speakers may deliberately violate these principles to express meanings beyond literal words, and that listeners, in turn, recognize such violations to draw pragmatic inferences. Building on Surian et al. (1996)'s study of children's sensitivity to violations of Gricean maxims, we introduce a novel benchmark to test whether language models pretrained on less than 10M and less than 100M tokens can distinguish maxim-adhering from maxim-violating utterances. We compare these BabyLMs across five maxims and situate their performance relative to children and a Large Language Model (LLM) pretrained on 3T tokens. We find that overall, models trained on less than 100M tokens outperform those trained on less than 10M, yet fall short of child-level and LLM competence. Our results suggest that modest data increases improve some aspects of pragmatic behavior, leading to finer-grained differentiation between pragmatic dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。