大模型会虚构人类无法感知的语法结构,且这些错误易被误认为正确。
LLMs Learn Constructions That Humans Do Not Know
- 通过上下文嵌入与提示探测,发现模型会生成人类不认可的语法构造。
- 模拟假设检验显示,虚假语法假设的准确率极高,易被误证真。
- 警示语言学研究需警惕模型引发的认知偏差,尤其在未知错误知识上。
本文研究了假阳性构造:大语言模型错误地将某些语法结构视为独立构造,但人类内省无法支持。通过基于上下文嵌入的行为探测任务和基于提示的元语言探测任务,可区分隐性与显性语言知识。两种方法均表明模型确实会虚构语法构造。进一步模拟假设检验发现,若语言学家错误假设这些构造存在,其高准确率将使其几乎必然被证实。这表明当前构造探测方法存在确认偏误,也引发对模型可能持有未知且错误句法知识的担忧。
原文摘要 · Abstract (English)
This paper investigates false positive constructions: grammatical structures which an LLM hallucinates as distinct constructions but which human introspection does not support. Both a behavioural probing task using contextual embeddings and a meta-linguistic probing task using prompts are included, allowing us to distinguish between implicit and explicit linguistic knowledge. Both methods reveal that models do indeed hallucinate constructions. We then simulate hypothesis testing to determine what would have happened if a linguist had falsely hypothesized that these hallucinated constructions do exist. The high accuracy obtained shows that such false hypotheses would have been overwhelmingly confirmed. This suggests that construction probing methods suffer from a confirmation bias and raises the issue of what unknown and incorrect syntactic knowledge these models also possess.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。