arXiv:2411.07533cs.CL2024-11ACL被引 4

用神经语言学方法揭示大模型形式与意义理解的差异

Large Language Models as Neurolinguistic Subjects: Discrepancy between Performance and Competence

  • 结合最小对立对与诊断探针分析模型各层激活模式
  • 发现模型表现力强但真实语言能力有限,形式理解优于意义
  • 提出中德文最小对立对数据集,支持跨语言研究

本研究通过区分心理语言学与神经语言学评估范式,探究大语言模型(LLMs)对符号形式(signifier)与意义(signified)的理解。传统心理语言学评估反映的是统计规律,未必体现真实语言能力。本文提出一种新神经语言学方法,结合最小对立对与诊断探针,分析模型各层激活模式,以深入考察其形式与意义表征的一致性。结果表明:(1)心理语言学与神经语言学方法显示,语言表现与能力存在显著差异;(2)直接概率测量无法准确评估语言能力;(3)指令微调可提升表现,但对能力影响甚微;(4)模型在形式理解上优于意义理解。此外,本文构建了中文(COMPS-ZH)和德文(COMPS-DE)最小对立对数据集,补充现有英文数据集。

原文摘要 · Abstract (English)

This study investigates the linguistic understanding of Large Language Models (LLMs) regarding signifier (form) and signified (meaning) by distinguishing two LLM assessment paradigms: psycholinguistic and neurolinguistic. Traditional psycholinguistic evaluations often reflect statistical rules that may not accurately represent LLMs' true linguistic competence. We introduce a neurolinguistic approach, utilizing a novel method that combines minimal pair and diagnostic probing to analyze activation patterns across model layers. This method allows for a detailed examination of how LLMs represent form and meaning, and whether these representations are consistent across languages. We found: (1) Psycholinguistic and neurolinguistic methods reveal that language performance and competence are distinct; (2) Direct probability measurement may not accurately assess linguistic competence; (3) Instruction tuning won't change much competence but improve performance; (4) LLMs exhibit higher competence and performance in form compared to meaning. Additionally, we introduce new conceptual minimal pair datasets for Chinese (COMPS-ZH) and German (COMPS-DE), complementing existing English datasets.

大模型理解神经语言学语言能力多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。