arXiv:2412.11318cs.CLcs.AI2024-12中稿 · CoLing 2025被引 2

用语言模型研究通用句的隐含量化,发现其高度依赖语境且多为弱普遍性。

Generics are puzzling. Can language models find the missing piece?

  • 构建上下文通用句数据集ConGen,分析语言模型对通用句的理解机制。
  • 20%的通用句表达弱普遍性,且比限定词更依赖具体语境。
  • 揭示语言模型中人类刻板印象的潜在偏见,适合关注自然语言语义的研究者。

通用句在不显式量化的情况下表达对世界的普遍性判断。尽管在日常交流中至关重要,但构建精确的语义框架仍具挑战,部分原因在于说话者使用通用句概括具有不同统计普遍性的属性。本文通过语言模型作为语言表征工具,研究通用句的隐含量化与语境敏感性。我们构建了包含2873个自然发生通用句与量化句的ConGen数据集,并提出基于意外度(surprisal)的p-acceptability指标,用于捕捉量化敏感性。实验表明,通用句的语境敏感性高于限定词,且约20%的自然通用句表达的是弱普遍性。此外,我们还探讨了人类刻板印象如何在语言模型中体现。

原文摘要 · Abstract (English)

Generic sentences express generalisations about the world without explicit quantification. Although generics are central to everyday communication, building a precise semantic framework has proven difficult, in part because speakers use generics to generalise properties with widely different statistical prevalence. In this work, we study the implicit quantification and context-sensitivity of generics by leveraging language models as models of language. We create ConGen, a dataset of 2873 naturally occurring generic and quantified sentences in context, and define p-acceptability, a metric based on surprisal that is sensitive to quantification. Our experiments show generics are more context-sensitive than determiner quantifiers and about 20% of naturally occurring generics we analyze express weak generalisations. We also explore how human biases in stereotypes can be observed in language models.

通用句语言模型语义理解偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。