arXiv:2410.16139cs.CL2024-10EMNLP被引 7

测试大模型对动作角色的敏感度,发现它与人类处理方式不同。

A Psycholinguistic Evaluation of Language Models' Sensitivity to Argument Roles

  • 用人类心理语言学实验方法评估模型对动词角色的判断能力
  • 模型能区分合理与不合理语境中的动词,但模式不同
  • 适合研究语言模型认知机制差异的学者参考

我们通过复现人类心理语言学研究,系统评估了大语言模型对论元角色(即谁对谁做了什么)的敏感性。在三个实验中,模型能够区分动词在合理与不合理语境中的表现,其中合理性由动词与其前置论元之间的关系决定。然而,所有模型均未表现出人类在实时动词预测中所展现的选择性模式。这表明,模型检测动词合理性的能力并非源于与人类相同的实时句法处理机制。

原文摘要 · Abstract (English)

We present a systematic evaluation of large language models' sensitivity to argument roles, i.e., who did what to whom, by replicating psycholinguistic studies on human argument role processing. In three experiments, we find that language models are able to distinguish verbs that appear in plausible and implausible contexts, where plausibility is determined through the relation between the verb and its preceding arguments. However, none of the models capture the same selective patterns that human comprehenders exhibit during real-time verb prediction. This indicates that language models' capacity to detect verb plausibility does not arise from the same mechanism that underlies human real-time sentence processing.

语言模型心理语言学论元角色

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。