用大模型预测词汇心理语言学特征,提升研究效率与准确性
Adding LLMs to the psycholinguistic norming toolbox: A practical guide to getting the most out of human ratings
- 通过基础模型和微调两种方式预测词汇特征,提升数据生成能力
- 基础模型相关性达0.8,微调后提升至0.9,接近人工标注水平
- 提供可复现框架与验证方法,适合心理学与语言学研究者使用
词汇级心理语言学规范为语言处理理论提供了实证支持,但获取基于人类的测量值往往不切实际或复杂。一种有前景的方法是利用大语言模型(LLMs)直接预测这些特征,该做法在心理语言学与认知科学中正迅速普及。然而,该方法的新颖性及大模型的相对不可解释性,要求研究者采用严谨方法论,指导实践、呈现多种路径,并揭示潜在局限性,某些情况下可能使大模型应用变得不切实际。本文提出一套全面的大模型估算词汇特征的方法,结合实践经验与实用建议。涵盖基础模型直接使用与模型微调两种策略,后者在特定场景下可显著提升性能。方法强调以人工“金标准”规范验证大模型生成数据。同时提供一个软件框架,支持商业与开源模型。案例研究以英语词汇熟悉度为例,基础模型达到0.8的斯皮尔曼相关性,微调后提升至0.9。该方法、框架与最佳实践旨在为未来基于大模型的心理语言学与词汇研究提供参考。
原文摘要 · Abstract (English)
Word-level psycholinguistic norms lend empirical support to theories of language processing. However, obtaining such human-based measures is not always feasible or straightforward. One promising approach is to augment human norming datasets by using Large Language Models (LLMs) to predict these characteristics directly, a practice that is rapidly gaining popularity in psycholinguistics and cognitive science. However, the novelty of this approach (and the relative inscrutability of LLMs) necessitates the adoption of rigorous methodologies that guide researchers through this process, present the range of possible approaches, and clarify limitations that are not immediately apparent, but may, in some cases, render the use of LLMs impractical. In this work, we present a comprehensive methodology for estimating word characteristics with LLMs, enriched with practical advice and lessons learned from our own experience. Our approach covers both the direct use of base LLMs and the fine-tuning of models, an alternative that can yield substantial performance gains in certain scenarios. A major emphasis in the guide is the validation of LLM-generated data with human "gold standard" norms. We also present a software framework that implements our methodology and supports both commercial and open-weight models. We illustrate the proposed approach with a case study on estimating word familiarity in English. Using base models, we achieved a Spearman correlation of 0.8 with human ratings, which increased to 0.9 when employing fine-tuned models. This methodology, framework, and set of best practices aim to serve as a reference for future research on leveraging LLMs for psycholinguistic and lexical studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。