改进实验刺激质量后,GPT-2在句法预测上表现显著提升。
Evaluating The Impact of Stimulus Quality in Investigations of LLM Language Performance
- 用语言学模板生成去噪刺激数据,减少歧义和复杂性干扰。
- 在新数据集上GPT-2的句法预测准确率明显高于旧基准。
- 研究提示:模型性能评估需重视刺激材料的质量设计。
近期利用大语言模型(LLMs)检验贫乏刺激论(APS)的研究,在不同句法现象上得出矛盾结论。本文提出假设:当前研究中使用的刺激材料存在词汇歧义和结构复杂性等特征,可能干扰模型表现。为此,本文以GPT-2为例,提出重评句法预测能力的方法:1)在先前使用过的过滤与未过滤刺激上建立基线;2)借助先进的生成模型Gemini 2.5 Pro Preview,基于语言学引导模板生成新的、优化后的刺激数据集(PG)。初步结果表明,相较于基线,GPT-2在新数据集上的表现显著提升,说明刺激质量对基于意外度(surprisal)的模型句法能力评估具有重要影响。
原文摘要 · Abstract (English)
Recent studies employing Large Language Models (LLMs) to test the Argument from the Poverty of the Stimulus (APS) have yielded contrasting results across syntactic phenomena. This paper investigates the hypothesis that characteristics of the stimuli used in recent studies, including lexical ambiguities and structural complexities, may confound model performance. A methodology is proposed for re-evaluating LLM competence on syntactic prediction, focusing on GPT-2. This involves: 1) establishing a baseline on previously used (both filtered and unfiltered) stimuli, and 2) generating a new, refined dataset using a state-of-the-art (SOTA) generative LLM (Gemini 2.5 Pro Preview) guided by linguistically-informed templates designed to mitigate identified confounds. Our preliminary findings indicate that GPT-2 demonstrates notably improved performance on these refined PG stimuli compared to baselines, suggesting that stimulus quality significantly influences outcomes in surprisal-based evaluations of LLM syntactic competency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。