arXiv:2512.12444cs.CL2025-12被引 1

GPT可有效替代人类评分,尤其在复杂隐喻的语义评估中表现可靠。

Can GPT replace human raters? Validity and reliability of machine-generated norms for metaphors

  • 用GPT生成687个英意双语隐喻的熟悉度、可理解性与意象性评分
  • 机器评分与人类评分高度相关,且能准确预测反应时与脑电波数据
  • 大模型表现更优,但对多模态和常规性隐喻的判断仍存偏差

随着大语言模型(LLMs)在科学研究中的广泛应用,其可信度问题日益重要。在心理语言学领域,尽管已有研究利用LLM自动补充单字评级数据并取得良好效果,但对复杂语义单位——如隐喻——的评级性能尚不清楚。本文首次系统评估了三个GPT模型对687个来自意大利修辞档案及三项英语研究的隐喻,在熟悉度、可理解性和意象性三方面生成的评分之有效性与可靠性。通过与人类数据对比,并检验其对行为反应和脑电生理信号的预测能力,发现机器评分与人工评分呈正相关:英语和意大利语隐喻的熟悉度达中至强相关,高感官运动负荷的隐喻相关性下降;意语意象性相关性更强。英语可理解性相关性最强。总体上,大模型优于小模型,且在熟悉度与意象性上人机差异更大。机器评分可显著预测反应时间与脑电幅值,效果接近人类评分。跨会话稳定性高。结论表明,尤其大模型可有效替代或补充人类进行隐喻属性评分,但在处理隐喻的常规性与多模态意义时仍需谨慎。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) are increasingly being used in scientific research, the issue of their trustworthiness becomes crucial. In psycholinguistics, LLMs have been recently employed in automatically augmenting human-rated datasets, with promising results obtained by generating ratings for single words. Yet, performance for ratings of complex items, i.e., metaphors, is still unexplored. Here, we present the first assessment of the validity and reliability of ratings of metaphors on familiarity, comprehensibility, and imageability, generated by three GPT models for a total of 687 items gathered from the Italian Figurative Archive and three English studies. We performed a thorough validation in terms of both alignment with human data and ability to predict behavioral and electrophysiological responses. We found that machine-generated ratings positively correlated with human-generated ones. Familiarity ratings reached moderate-to-strong correlations for both English and Italian metaphors, although correlations weakened for metaphors with high sensorimotor load. Imageability showed moderate correlations in English and moderate-to-strong in Italian. Comprehensibility for English metaphors exhibited the strongest correlations. Overall, larger models outperformed smaller ones and greater human-model misalignment emerged with familiarity and imageability. Machine-generated ratings significantly predicted response times and the EEG amplitude, with a strength comparable to human ratings. Moreover, GPT ratings obtained across independent sessions were highly stable. We conclude that GPT, especially larger models, can validly and reliably replace - or augment - human subjects in rating metaphor properties. Yet, LLMs align worse with humans when dealing with conventionality and multimodal aspects of metaphorical meaning, calling for careful consideration of the nature of stimuli.

大模型隐喻心理语言学评分替代

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。