用大模型预测句子记忆难易和阅读时长,效果优于传统方法。
To Words and Beyond: Probing Large Language Models for Sentence-Level Psycholinguistic Norms of Memorability and Reading Times
- 通过微调让大模型预测句子级心理语言学特征
- 预测结果与人类数据高度相关,超越基础模型表现
- 零样本/少样本效果不稳,需谨慎使用大模型作认知代理
大语言模型(LLMs)已被证明能以零样本方式生成词语或短语的心理语言学特征(如情感、唤醒度、具体性)估计值,且与人类判断高度相关。但对于词汇决策时间或习得年龄等特征,通常需要监督微调才能达到真实值水平。本文将该方法拓展至此前未研究的句子级特征——句子记忆难度与阅读时间,这些特征依赖于句内多个词之间的语义关系。实验表明,经微调后,模型对这些特征的估计与人类基准数据高度相关,并显著优于可解释基线模型。但零样本和少量样本下的表现参差不齐,进一步说明在将大模型提示作为人类认知指标代理时需格外谨慎。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have recently been shown to produce estimates of psycholinguistic norms, such as valence, arousal, or concreteness, for words and multiword expressions, that correlate with human judgments. These estimates are obtained by prompting an LLM, in zero-shot fashion, with a question similar to those used in human studies. Meanwhile, for other norms such as lexical decision time or age of acquisition, LLMs require supervised fine-tuning to obtain results that align with ground-truth values. In this paper, we extend this approach to the previously unstudied features of sentence memorability and reading times, which involve the relationship between multiple words in a sentence-level context. Our results show that via fine-tuning, models can provide estimates that correlate with human-derived norms and exceed the predictive power of interpretable baseline predictors, demonstrating that LLMs contain useful information about sentence-level features. At the same time, our results show very mixed zero-shot and few-shot performance, providing further evidence that care is needed when using LLM-prompting as a proxy for human cognitive measures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。