用词嵌入分析马来语词素透明度,发现语义中心距能最好预测阅读反应时间。
An Exploratory Analysis on the Explanatory Potential of Embedding-Based Measures of Semantic Transparency for Malay Word Recognition
- 基于词嵌入计算词素透明度,通过语义空间几何结构评估复杂词
- 5种透明度度量均显著预测词汇判断反应时,其中词与语义中心距离效果最佳
- 适合研究语言认知、词法加工的学者,尤其关注语义表示与阅读行为关系
词汇形态加工研究表明语义透明度对单词识别至关重要,但其计算操作化仍存争议。本文旨在探索基于词嵌入的语义透明度度量及其对阅读的影响。首先,对4,226个马来语带前缀词进行t-SNE聚类分析,观察到不同前缀类别的复杂词在语义空间中形成多个聚类。随后构建五种简单度量,并检验其是否为词汇判断反应时的显著预测因子。通过两组线性判别分析,分别以词嵌入或移位向量(基词与派生词向量差)预测词前缀,模型预测准确率反映前缀透明度。进一步计算三类度量:词与同前缀词语义中心的相似性、词与基词移位向量的相似性、词与功能词素组合语义空间模型预测词的相似性。在一系列广义加性混合模型中,所有度量在控制词频、词长和词族大小后均能显著预测反应时,其中词与语义中心相关性作为预测变量的模型拟合最优。
原文摘要 · Abstract (English)
Studies of morphological processing have shown that semantic transparency is crucial for word recognition. Its computational operationalization is still under discussion. Our primary objectives are to explore embedding-based measures of semantic transparency, and assess their impact on reading. First, we explored the geometry of complex words in semantic space. To do so, we conducted a t-distributed Stochastic Neighbor Embedding clustering analysis on 4,226 Malay prefixed words. Several clusters were observed for complex words varied by their prefix class. Then, we derived five simple measures, and investigated whether they were significant predictors of lexical decision latencies. Two sets of Linear Discriminant Analyses were run in which the prefix of a word is predicted from either word embeddings or shift vectors (i.e., a vector subtraction of the base word from the derived word). The accuracy with which the model predicts the prefix of a word indicates the degree of transparency of the prefix. Three further measures were obtained by comparing embeddings between each word and all other words containing the same prefix (i.e., centroid), between each word and the shift from their base word, and between each word and the predicted word of the Functional Representations of Affixes in Compositional Semantic Space model. In a series of Generalized Additive Mixed Models, all measures predicted decision latencies after accounting for word frequency, word length, and morphological family size. The model that included the correlation between each word and their centroid as a predictor provided the best fit to the data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。