arXiv:2409.15890cs.CL2024-09被引 15

用心理语言学实验评估20个大模型的语言人类相似度

HLB: Benchmarking LLMs' Humanlikeness in Language Use

  • 设计10项心理语言学实验,从语音到语篇全面测试语言模型
  • 对比2000+真人与模型的响应分布,量化人类相似度
  • 发现性能提升未必带来更像人,为评估模型提供新框架

随着合成数据在语言模型训练中日益普遍,尤其是生成对话数据,人们担忧这些模型可能偏离真实的人类语言模式,从而丧失人际交流中的丰富性与创造性。这凸显了评估语言模型在真实语言使用中人类相似度的紧迫性。本文提出一个全面的人类相似度基准(HLB),通过10项心理语言学实验,评估20个大语言模型(LLMs)在语音、词汇、句法、语义和语篇等核心语言层面的表现。为建立参照标准,我们收集了超过2000名人类参与者的数据,并将其与模型输出进行对比。为实现严谨评估,我们开发了一种编码算法,可准确识别语言使用模式,提取各任务的响应分布。通过比较人类与模型的响应分布,我们以分布相似性量化人类相似度。结果揭示了不同语言层次上模型复现人类响应的细微差异。值得注意的是,其他性能指标的提升并不必然带来更高的人类相似度,甚至可能导致下降。该基准首次系统性引入心理语言学方法,为评估大模型语言使用的自然程度提供了新范式。

原文摘要 · Abstract (English)

As synthetic data becomes increasingly prevalent in training language models, particularly through generated dialogue, concerns have emerged that these models may deviate from authentic human language patterns, potentially losing the richness and creativity inherent in human communication. This highlights the critical need to assess the humanlikeness of language models in real-world language use. In this paper, we present a comprehensive humanlikeness benchmark (HLB) evaluating 20 large language models (LLMs) using 10 psycholinguistic experiments designed to probe core linguistic aspects, including sound, word, syntax, semantics, and discourse (see https://huggingface.co/spaces/XufengDuan/HumanLikeness). To anchor these comparisons, we collected responses from over 2,000 human participants and compared them to outputs from the LLMs in these experiments. For rigorous evaluation, we developed a coding algorithm that accurately identified language use patterns, enabling the extraction of response distributions for each task. By comparing the response distributions between human participants and LLMs, we quantified humanlikeness through distributional similarity. Our results reveal fine-grained differences in how well LLMs replicate human responses across various linguistic levels. Importantly, we found that improvements in other performance metrics did not necessarily lead to greater humanlikeness, and in some cases, even resulted in a decline. By introducing psycholinguistic methods to model evaluation, this benchmark offers the first framework for systematically assessing the humanlikeness of LLMs in language use.

大模型评估人类相似度心理语言学基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。