用语言模型嵌入分析阅读时的语义关联,发现句向量更可靠。
Modeling semantic association in self-paced reading with language model embeddings

- 用不同语言模型和上下文长度计算语义关联
- 只有句向量能稳定预测脑电与阅读时间
- 方法选择影响结果,建议用句向量
语义关联是阅读理解的重要因素,即使考虑了词可预测性。近期研究显示语言模型(LM)嵌入可用于量化语义关联,但其操作方式多样。本研究基于荷兰语自然文本的联合脑电图(EEG)与自定速阅读数据,使用十种不同实现方式计算语义关联,包括不同嵌入模型和上下文长度。通过贝叶斯分层模型和贝叶斯因子分析其对N400成分及自定速阅读时间的影响。结果显示,嵌入模型的选择会改变语义关联对神经与行为指标的影响估计;唯有依赖句向量的实现方式,在神经与行为测量上均表现出超越词可预测性的可靠语义关联信号。研究强调了量化语义关联时方法选择的重要性。
原文摘要 · Abstract (English)
Semantic association between a word and its context has been identified as an important component of reading comprehension, even when word predictability is accounted for. Recent research has highlighted the potential of language model ( LM) embeddings to quantify semantic association. Yet, embedding-based semantic association have been operationalized in a myriad of ways. In this study, we use embeddings from LMs to estimate semantic association on a corpus of joint electroencephalography (EEG) and self-paced reading of natural, Dutch texts. Semantic association is calculated in ten different implementations that vary the embedding model and context lengths. The effects of semantic association across the different implementations on the N400 and self-paced reading times are examined using Bayesian hierarchical models and Bayes factor. The results show that the choice of embedding model can alter the estimated effect of semantic association on both the N400 and self-paced reading times. Furthermore, the results demonstrate a promising potential of sentence embeddings for capturing semantic association, as only implementations relying on sentence embeddings indicate reliable results of semantic association beyond word predictability on both neural and behavioral measures. Together, these findings highlight the importance of methodological choices in quantifying semantic association.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。