用直接最小对分析法发现大模型能掌握复杂句法结构。
Exploring Gaps in the APS: Direct Minimal Pair Analysis in LLM Syntactic Assessments
- 采用直接最小对对比法,系统检验模型对填空-缺口依赖的掌握程度。
- GPT-2在4种测试条件下均表现良好,说明其具备填空缺口句法能力。
- 该方法比差分差分法更透明,适合评估大模型的句法理解力。
近期研究利用大语言模型(LLMs)通过基于意外度的指标检验复杂句法的可学习性,但结论分歧引发对其有效性质疑。Wilcox等(2024)通过直接最小对比较('wh-effect')证明模型能成功泛化填空-缺口依赖知识;而Lan等(2024)使用差分差分(DiD)指标发现模型在寄生缺口(PGs)任务中表现不佳。本文主张直接最小对方法具有更高诊断透明性。我们生成了完整的8种排列组合的精炼寄生缺口刺激材料,并对前文使用的GPT-2模型进行系统性的威尔科克斯风格wh-effect分析。结果显示,GPT-2在所有四个测试条件下均表现成功,表明其在复杂寄生缺口环境中仍具备填空-缺口许可原则的稳健知识。这一结果与基于DiD指标的模糊结论形成对比,凸显评估指标选择对判断大模型句法能力的关键影响。
原文摘要 · Abstract (English)
Recent studies probing the Argument from the Poverty of the Stimulus (APS) have applied Large Language Models (LLMs) to test the learnability of complex syntax through surprisal-based metrics. However, divergent conclusions raise questions concerning the insights these metrics offer. While Wilcox et al. (2024) used direct minimal pair comparisons (the "wh-effect") to demonstrate that models successfully generalise knowledge of filler-gap dependencies, Lan et al. (2024) used a Difference-in-Differences (DiD) metric and found that models largely fail on parasitic gaps (PGs). This paper argues that the direct minimal pair approach offers greater diagnostic transparency. We demonstrate this by generating a full 8-permutation paradigm of refined PG stimuli and evaluating the GPT-2 model used in previous studies with a systematic Wilcox-style wh-effect analysis. Our results show that GPT-2 succeeds across all four tested conditions, indicating robust knowledge of filler-gap licensing principles even in complex PG environments. This finding, which contrasts with the more ambiguous results from DiD-style metrics, suggests that the choice of evaluation metric is critical for assessing an LLM's syntactic competence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。