用贝叶斯推理解释语言错觉强度差异,模型预测精准且涵盖新现象。
Graded strength of comparative illusions is explained by Bayesian inference
- 结合语言模型与人类行为数据,构建贝叶斯后验概率模型预测错觉强度。
- 模型准确解释了比较错觉的强弱梯度及代词与名词短语的差异效应。
- 适合对语言认知机制、计算语言学感兴趣的读者。
与视觉处理类似,语言理解也易受错觉影响,如比较错觉(CI),例如“去俄罗斯的学生比我还多”这类句子在语义上不成立,但听者仍认为可接受。已有研究提出,这可用噪声信道中的贝叶斯推断解释:句子的理解后验概率与其先验概率及被误传为该句的可能性成正比。本研究在此基础上,通过融合统计语言模型与人类行为数据,构建了一个定量预测模型,能精确解释比较错觉的细粒度强度变化。该模型还首次成功解释了由代词或完整名词短语作从句主语导致的未解效应。结果支持噪声信道理论作为统一的计算层级模型,能解释多种语言处理现象,包括有错觉与无错觉情境。
原文摘要 · Abstract (English)
Like visual processing, language processing is susceptible to illusions in which people systematically misperceive stimuli. In one such case--the comparative illusion (CI), e.g., More students have been to Russia than I have--comprehenders tend to judge the sentence as acceptable despite its underlying nonsensical comparison. Prior research has argued that this phenomenon can be explained as Bayesian inference over a noisy channel: the posterior probability of an interpretation of a sentence is proportional to both the prior probability of that interpretation and the likelihood of corruption into the observed (CI) sentence. Initial behavioral work has supported this claim by evaluating a narrow set of alternative interpretations of CI sentences and showing that comprehenders favor interpretations that are more likely to have been corrupted into the illusory sentence. In this study, we replicate and go substantially beyond this earlier work by directly predicting the strength of illusion with a quantitative model of the posterior probability of plausible interpretations, which we derive through a novel synthesis of statistical language models with human behavioral data. Our model explains not only the fine gradations in the strength of CI effects, but also a previously unexplained effect caused by pronominal vs. full noun phrase than-clause subjects. These findings support a noisy-channel theory of sentence comprehension by demonstrating that the theory makes novel predictions about the comparative illusion that bear out empirically. This outcome joins related evidence of noisy channel processing in both illusory and non-illusory contexts to support noisy channel inference as a unified computational-level theory of diverse language processing phenomena.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。