用心理语言学特征评估大模型与人类认知的对齐程度
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans
- 用心理学实验数据对比大模型与人类对词汇的情感、感官等特征的感知
- 格拉斯哥数据集对齐度高于兰卡斯特数据集,尤其在情绪和具体性上表现更好
- 揭示大模型缺乏具身认知,适合研究人机认知差异的学者参考
大模型评估长期聚焦于推理、问答等任务的客观性能,但许多语言特征难以量化。心理语言学通过大规模人类实验,为数千个词语建立了包括唤醒度、效价、具体性、感官关联等十三种特征的评分体系。本文利用格拉斯哥和兰卡斯特两个语料库,评估代表性大模型在这些特征上的对齐程度。结果显示,模型在格拉斯哥数据集(涵盖唤醒度、效价、支配感、具体性、意象性、熟悉度及性别)上的对齐度普遍较高;而在兰卡斯特数据集(涵盖内省、味觉、嗅觉、触觉、听觉和视觉)上的对齐度较低。这表明当前大模型在模拟人类感官联想方面存在局限,可能源于缺乏具身认知,凸显了使用心理语言学数据评估大模型的重要性。
原文摘要 · Abstract (English)
The evaluation of LLMs has so far focused primarily on how well they can perform different tasks such as reasoning, question-answering, paraphrasing, or translating. For most of these tasks, performance can be measured with objective metrics, such as the number of correct answers. However, other language features are not easily quantified. For example, arousal, concreteness, or gender associated with a given word, as well as the extent to which we experience words with senses and relate them to a specific sense. Those features have been studied for many years by psycholinguistics, conducting large-scale experiments with humans to produce ratings for thousands of words. This opens an opportunity to evaluate how well LLMs align with human ratings on these word features, taking advantage of existing studies that cover many different language features in a large number of words. In this paper, we evaluate the alignment of a representative group of LLMs with human ratings on two psycholinguistic datasets: the Glasgow and Lancaster norms. These datasets cover thirteen features over thousands of words. The results show that alignment is \textcolor{black}{generally} better in the Glasgow norms evaluated (arousal, valence, dominance, concreteness, imageability, familiarity, and gender) than on the Lancaster norms evaluated (introceptive, gustatory, olfactory, haptic, auditory, and visual). This suggests a potential limitation of current LLMs in aligning with human sensory associations for words, which may be due to their lack of embodied cognition present in humans and illustrates the usefulness of evaluating LLMs with psycholinguistic datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。