用贝叶斯滤波提升语言模型对歧义的不确定性建模能力
Ensemble Kalman filter for uncertainty in human language comprehension
- 将句子理解建模为贝叶斯逆问题,引入集合卡尔曼滤波量化不确定性
- 在句法语义反转句上,贝叶斯方法比最大似然估计更贴近人类认知表现
- 适合研究语言认知、不确定性建模与神经网络可解释性的学者
人工神经网络(ANN)广泛用于句子处理建模,但通常呈现确定性行为,与人类在面对模糊或意外输入时能管理不确定性的认知方式形成对比。这一差异在反向异常句(含意外角色反转的句子)中尤为明显,挑战了句法和语义理解,暴露出传统ANN模型(如句段模型,SG模型)的局限性。为解决这些问题,我们提出一种基于贝叶斯框架的句子理解方法,采用集合卡尔曼滤波(EnKF)扩展实现贝叶斯推断,以量化不确定性。通过将语言理解视为贝叶斯逆问题,该方法显著增强了SG模型对人类句子处理中不确定性表征的能力。数值实验及与最大似然估计(MLE)的对比表明,贝叶斯方法在处理语言歧义时,能更准确地逼近人类认知过程。
原文摘要 · Abstract (English)
Artificial neural networks (ANNs) are widely used in modeling sentence processing but often exhibit deterministic behavior, contrasting with human sentence comprehension, which manages uncertainty during ambiguous or unexpected inputs. This is exemplified by reversal anomalies-sentences with unexpected role reversals that challenge syntax and semantics-highlighting the limitations of traditional ANN models, such as the Sentence Gestalt (SG) Model. To address these limitations, we propose a Bayesian framework for sentence comprehension, applying an extension of the ensemble Kalman filter (EnKF) for Bayesian inference to quantify uncertainty. By framing language comprehension as a Bayesian inverse problem, this approach enhances the SG model's ability to reflect human sentence processing with respect to the representation of uncertainty. Numerical experiments and comparisons with maximum likelihood estimation (MLE) demonstrate that Bayesian methods improve uncertainty representation, enabling the model to better approximate human cognitive processing when dealing with linguistic ambiguities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。