用词嵌入预测感官运动规范,揭示语言与身体经验的关联
Words that make SENSE: Sensorimotor Norms in Learned Lexical Token Representations
- 构建可从词向量推断感官运动特性的学习模型
- 人类对新词的选择率与模型评分在6个感知维度显著相关
- 发现内感受规范存在系统性语音象似模式,可生成候选音义组合
尽管词嵌入依赖共现模式获得语义,但人类语言理解根植于感官与运动体验。本文提出 $ ext{SENSE}$(感官运动嵌入规范评分引擎),一种从词法嵌入预测兰卡斯特感官运动规范的学习投影模型。我们还开展了一项行为实验,281名参与者从候选陌生词中选择能引发特定感官运动联想的词,结果显示人类选择率与 $ ext{SENSE}$ 评分在11个模态中的6个维度上存在统计显著相关性。对这些陌生词的选择率进行子词分析,揭示了内感受规范存在系统性语音象似模式,为从文本数据中计算生成候选音义特征提供了路径。
原文摘要 · Abstract (English)
While word embeddings derive meaning from co-occurrence patterns, human language understanding is grounded in sensory and motor experience. We present $\text{SENSE}$ $(\textbf{S}\text{ensorimotor }$ $\textbf{E}\text{mbedding }$ $\textbf{N}\text{orm }$ $\textbf{S}\text{coring }$ $\textbf{E}\text{ngine})$, a learned projection model that predicts Lancaster sensorimotor norms from word lexical embeddings. We also conducted a behavioral study where 281 participants selected which among candidate nonce words evoked specific sensorimotor associations, finding statistically significant correlations between human selection rates and $\text{SENSE}$ ratings across 6 of the 11 modalities. Sublexical analysis of these nonce words selection rates revealed systematic phonosthemic patterns for the interoceptive norm, suggesting a path towards computationally proposing candidate phonosthemes from text data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。