用预训练模型分析词语分布,发现语法结构可被统计关联揭示。
Constructions are Revealed in Word Distributions
- 将RoBERTa模型视为语言分布代理,通过统计关联挖掘构造
- 成功识别出语义不同但形式相似的构造及抽象填空的框架构造
- 提示统计关联是学习者的重要线索,但非唯一依据
构式语法认为构式(形式-意义配对)通过语言经验习得(分布学习假说)。但文本中实际包含多少构式信息?现有语料分析提供部分答案,但无法回答“什么导致了某个词出现”这类反事实问题。这需要可计算的语言字符串分布模型——即预训练语言模型(PLMs)。本文将RoBERTa模型作为该分布的代理,假设构式会以统计亲和性模式在其中显现。实验支持此假设:许多构式可被稳健区分,包括(i)语义不同但表面相似的困难案例,以及(ii)可由抽象词类填充的框架构造。尽管如此,我们也提供了定性证据表明,仅靠统计亲和性可能不足以从文本中识别所有构式。因此,统计亲和性可能是学习者可用的重要但不完整的信号。
原文摘要 · Abstract (English)
Construction grammar posits that constructions, or form-meaning pairings, are acquired through experience with language (the distributional learning hypothesis). But how much information about constructions does this distribution actually contain? Corpus-based analyses provide some answers, but text alone cannot answer counterfactual questions about what \emph{caused} a particular word to occur. This requires computable models of the distribution over strings -- namely, pretrained language models (PLMs). Here, we treat a RoBERTa model as a proxy for this distribution and hypothesize that constructions will be revealed within it as patterns of statistical affinity. We support this hypothesis experimentally: many constructions are robustly distinguished, including (i) hard cases where semantically distinct constructions are superficially similar, as well as (ii) \emph{schematic} constructions, whose ``slots'' can be filled by abstract word classes. Despite this success, we also provide qualitative evidence that statistical affinity alone may be insufficient to identify all constructions from text. Thus, statistical affinity is likely an important, but partial, signal available to learners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。