arXiv:2507.10342cs.CL2025-07

用大模型复现人类语言实验,发现AI与人判断高度一致。

Using AI to replicate human experimental results: a motion study

  • 用GPT-4复刻四项人类语言实验,测试情感意义随动作动词的变化。
  • 四组实验中,人类与AI评分相关性达0.73至0.96,整体模式一致。
  • 适合语言学研究者探索假设、扩展数据,提升实验规模与效率。

本文探讨大型语言模型(LLMs)作为语言学研究中可靠分析工具的潜力,聚焦于动作方式动词在时间表达中引发的情感意义演变。尽管GPT-4等模型在多项任务中表现优异,其复现人类细微判断的能力仍存疑。我们先后对人类参与者和一个LLM开展四项心理语言学实验:关于涌现意义、情感极性变化、情绪语境下的动词选择,以及句子与表情符号的关联。所有实验结果均显示人类与AI响应高度吻合,统计分析(如斯皮尔曼等级相关系数ρ = .73–.96)表明评分模式与分类选择具有强相关性。尽管个别情况下存在微小差异,但未改变整体解释结论。这些发现为使用LLMs辅助传统人类实验提供了有力证据,使其能在不牺牲解释有效性的情况下实现更大规模研究。这一一致性不仅强化了以往基于人类研究的实证基础,还为通过人工智能进行假说生成与数据扩展开辟了新路径。最终,本研究支持将LLMs视为语言探究中可信且富有信息量的合作者。

原文摘要 · Abstract (English)

This paper explores the potential of large language models (LLMs) as reliable analytical tools in linguistic research, focusing on the emergence of affective meanings in temporal expressions involving manner-of-motion verbs. While LLMs like GPT-4 have shown promise across a range of tasks, their ability to replicate nuanced human judgements remains under scrutiny. We conducted four psycholinguistic studies (on emergent meanings, valence shifts, verb choice in emotional contexts, and sentence-emoji associations) first with human participants and then replicated the same tasks using an LLM. Results across all studies show a striking convergence between human and AI responses, with statistical analyses (e.g., Spearman's rho = .73-.96) indicating strong correlations in both rating patterns and categorical choices. While minor divergences were observed in some cases, these did not alter the overall interpretative outcomes. These findings offer compelling evidence that LLMs can augment traditional human-based experimentation, enabling broader-scale studies without compromising interpretative validity. This convergence not only strengthens the empirical foundation of prior human-based findings but also opens possibilities for hypothesis generation and data expansion through AI. Ultimately, our study supports the use of LLMs as credible and informative collaborators in linguistic inquiry.

语言模型心理语言学实验复现大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。