用大模型替代人类做语言学实验,效果更优且可控。
Are Large Language Models the future crowd workers of Linguistics?
- 用GPT-4o-mini零样本提示复现两个语言学实验
- 大模型在关键任务上表现优于人类参与者
- 思维链提示可进一步接近人类表现,适合人文研究者参考
人工参与的数据采集是实证语言学研究的核心方法之一,但样本量从少量到众包规模不等。尽管数据丰富,这些方法存在注意力控制差、众包环境工作条件差、实验设计耗时等问题。为此,本文探讨大语言模型(LLMs)是否能克服上述障碍并融入语言学研究流程。通过复现Cruz(2023)和Lombard等(2021)的两个强制性语料获取任务,使用OpenAI的GPT-4o-mini模型在零样本提示下完成实验。结果表明,该模型在任务中表现出色且高度灵活,整体性能优于人类被试。第二项复制研究进一步发现,采用思维链(Chain-of-Thought, CoT)提示策略后,模型在关键项与填充项上的表现更接近人类水平。由于本研究规模有限,未来需进一步探索大模型在语言学及其他人文学科中的应用潜力。
原文摘要 · Abstract (English)
Data elicitation from human participants is one of the core data collection strategies used in empirical linguistic research. The amount of participants in such studies may vary considerably, ranging from a handful to crowdsourcing dimensions. Even if they provide resourceful extensive data, both of these settings come alongside many disadvantages, such as low control of participants' attention during task completion, precarious working conditions in crowdsourcing environments, and time-consuming experimental designs. For these reasons, this research aims to answer the question of whether Large Language Models (LLMs) may overcome those obstacles if included in empirical linguistic pipelines. Two reproduction case studies are conducted to gain clarity into this matter: Cruz (2023) and Lombard et al. (2021). The two forced elicitation tasks, originally designed for human participants, are reproduced in the proposed framework with the help of OpenAI's GPT-4o-mini model. Its performance with our zero-shot prompting baseline shows the effectiveness and high versatility of LLMs, that tend to outperform human informants in linguistic tasks. The findings of the second replication further highlight the need to explore additional prompting techniques, such as Chain-of-Thought (CoT) prompting, which, in a second follow-up experiment, demonstrates higher alignment to human performance on both critical and filler items. Given the limited scale of this study, it is worthwhile to further explore the performance of LLMs in empirical Linguistics and in other future applications in the humanities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。