arXiv:2505.16277cs.CL2025-05被引 3

用自发口语数据评估大模型认知合理性,发现语音训练更优。

Spontaneous Speech Variables for Evaluating LLMs Cognitive Plausibility

  • 用口语语料提取说话简化和语调突出等生成变量
  • 微调后模型预测准确率显著高于基线,语音训练效果更好
  • 适合关注大模型认知真实性的研究者参考

大型语言模型在自然语言处理中的成就,尤其在高资源语言上,亟需从认知角度深入理解其特性。已有研究通过测试模型对语言加工过程中行为(如眼动停留)和生理(如脑电反应)变量的预测能力来评估人工模型。本文提出利用自发口语语料库提取生产性变量(如说话简化、语调突出),并以此评估模型表现。具体地,我们从语料中提取这些变量,并测试经过标准训练流程、在不同预训练数据集(书面、口语、混合)上训练的模型对其预测能力。结果表明,经微调后,模型对这些生产变量的预测显著优于基线;且使用口语语料训练的模型表现优于纯书面语训练模型。该研究为利用高质量语音语料作为大模型评估基准提供了支持。

原文摘要 · Abstract (English)

The achievements of Large Language Models in Natural Language Processing, especially for high-resource languages, call for a better understanding of their characteristics from a cognitive perspective. Researchers have attempted to evaluate artificial models by testing their ability to predict behavioral (e.g., eye-tracking fixations) and physiological (e.g., brain responses) variables during language processing (e.g., reading/listening). In this paper, we propose using spontaneous speech corpora to derive production variables (speech reductions, prosodic prominences) and applying them in a similar fashion. More precisely, we extract. We then test models trained with a standard procedure on different pretraining datasets (written, spoken, and mixed genres) for their ability to predict these two variables. Our results show that, after some fine-tuning, the models can predict these production variables well above baselines. We also observe that spoken genre training data provides more accurate predictions than written genres. These results contribute to the broader effort of using high-quality speech corpora as benchmarks for LLMs.

大模型评估认知合理性口语语料生成变量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。