对比词向量与上下文嵌入对复合词语义的捕捉能力
The aftermath of compounds: Investigating Compounds and their Semantic Representations
- 用人类评分数据对比GloVe和BERT在复合词语义上的表现
- BERT在语义透明度上相关性更高,预测性是关键因素
- 适合关注语义建模与心理语言学交叉的研究者
本研究考察计算嵌入在英语复合词处理中与人类语义判断的对齐程度。通过对比静态词向量(GloVe)与上下文嵌入(BERT)与心理语言学数据中的人类评分(词素意义主导度LMD、语义透明度ST),利用爱丁堡联想词典、英国国家语料库(BNC)和LaDEC的关联强度、频率与可预测性指标,计算嵌入派生的LMD与ST,并通过斯皮尔曼相关与回归分析评估其与人类判断的关系。结果表明,BERT嵌入在捕捉组合语义方面优于GloVe,且可预测性在人类与模型数据中均为语义透明度的重要预测因子。研究推进了计算心理语言学的发展,揭示了复合词处理的关键驱动因素,并为基于嵌入的语义建模提供洞见。
原文摘要 · Abstract (English)
This study investigates how well computational embeddings align with human semantic judgments in the processing of English compound words. We compare static word vectors (GloVe) and contextualized embeddings (BERT) against human ratings of lexeme meaning dominance (LMD) and semantic transparency (ST) drawn from a psycholinguistic dataset. Using measures of association strength (Edinburgh Associative Thesaurus), frequency (BNC), and predictability (LaDEC), we compute embedding-derived LMD and ST metrics and assess their relationships with human judgments via Spearmans correlation and regression analyses. Our results show that BERT embeddings better capture compositional semantics than GloVe, and that predictability ratings are strong predictors of semantic transparency in both human and model data. These findings advance computational psycholinguistics by clarifying the factors that drive compound word processing and offering insights into embedding-based semantic modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。