人类与人工智能在语言结构处理上表现出神经表征的收敛性。
Convergent Representations of Linguistic Constructions in Human and Artificial Neural Systems
- 通过脑电实验验证语言模型对语法结构的表征可映射到人脑活动。
- 句末位置的α波段出现结构特异性神经信号,分类准确率高。
- 结果揭示生物与人工系统在语言抽象表示上的深层一致性。
理解大脑如何处理语言结构是认知神经科学和语言学的核心挑战。近期计算研究显示,人工神经语言模型会自发形成论元结构(ASC)的差异化表征,并预测其在处理过程中的出现时机与方式。本研究利用脑电图(EEG)在10名母语为英语的参与者中检验这些预测,让他们聆听200个合成句子,涵盖四种构造类型(及物、双宾、致使运动、结果性)。结合时频分析、特征提取与机器学习分类,发现构造特异性神经信号主要在句末位置出现,此时论元结构完全明确,且以α波段最为显著。成对分类显示,尤其在双宾与结果性构造间区分可靠,其他组合存在重叠。关键的是,这些效应的时间演变与相似性结构与循环和变压器语言模型中的模式一致,即构造表征在整合处理阶段产生。研究支持语言构造作为独立的形式-意义映射在神经层面被编码的观点,符合构式语法理论,并表明生物与人工系统在相似表征解决方案上出现收敛。更广泛而言,这种收敛支持学习系统在潜在表征空间中发现稳定区域——最近称为柏拉图表征空间——从而限制高效语言抽象的出现。
原文摘要 · Abstract (English)
Understanding how the brain processes linguistic constructions is a central challenge in cognitive neuroscience and linguistics. Recent computational studies show that artificial neural language models spontaneously develop differentiated representations of Argument Structure Constructions (ASCs), generating predictions about when and how construction-level information emerges during processing. The present study tests these predictions in human neural activity using electroencephalography (EEG). Ten native English speakers listened to 200 synthetically generated sentences across four construction types (transitive, ditransitive, caused-motion, resultative) while neural responses were recorded. Analyses using time-frequency methods, feature extraction, and machine learning classification revealed construction-specific neural signatures emerging primarily at sentence-final positions, where argument structure becomes fully disambiguated, and most prominently in the alpha band. Pairwise classification showed reliable differentiation, especially between ditransitive and resultative constructions, while other pairs overlapped. Crucially, the temporal emergence and similarity structure of these effects mirror patterns in recurrent and transformer-based language models, where constructional representations arise during integrative processing stages. These findings support the view that linguistic constructions are neurally encoded as distinct form-meaning mappings, in line with Construction Grammar, and suggest convergence between biological and artificial systems on similar representational solutions. More broadly, this convergence is consistent with the idea that learning systems discover stable regions within an underlying representational landscape - recently termed a Platonic representational space - that constrains the emergence of efficient linguistic abstractions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。