首次构建多模态句法激活数据集,验证视觉如何引导模型句法选择。
Towards Human Cognition: Visual Context Guides Syntactic Priming in Fusion-Encoded Models
- 设计融合编码架构,让视觉上下文影响句子结构生成
- 发现融合模型的句法保留与视觉相似度强相关
- 提出无参考评估指标,适合研究多模态语言模型认知机制
句法激活是一种认知现象:接触某种句法结构后,后续表达更可能重复该结构。尽管人类在多种语言情境中均表现出一致的句法激活效应,但多模态大语言模型(MLLMs)是否具备类似行为尚不明确。本文提出首个多模态句法激活数据集PRISMATIC,为研究句法-视觉交互提供标准化基准。设计了句法保留指数(SPI),一种专门用于评估句法激活的无参考指标。通过两种不同多模态编码架构的模型对比实验发现,两类模型均表现出可比的句法激活效应;但仅融合编码模型在句法激活强度与视觉相似度之间呈现显著正相关,暗示其更接近人类心理语言学模式。本研究为理解多模态语言模型中句法信息处理提供了新视角。
原文摘要 · Abstract (English)
Structural priming is a cognitive phenomenon where exposure to a particular syntactic structure increases the likelihood of producing the same structure in subsequent utterances. While humans consistently demonstrate structural priming effects across various linguistic contexts, it remains unclear whether multimodal large language models (MLLMs) exhibit similar syntactic preservation behaviors. We introduce PRISMATIC, the first multimodal structural priming dataset, which advances computational linguistics by providing a standardized benchmark for investigating syntax-vision interactions. We propose the Syntactic Preservation Index (SPI), a novel reference-free evaluation metric designed specifically to assess structural priming effects in sentence level. Using this metric, we constructed and tested models with two different multimodal encoding architectures to investigate their structural preservation capabilities. Our experimental results demonstrate that models with both encoding methods show comparable syntactic priming effects. However, only fusion-encoded models exhibit robust positive correlations between priming effects and visual similarity, suggesting a cognitive process more aligned with human psycholinguistic patterns. This work provides new insights into evaluating and understanding how syntactic information is processed in multimodal language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。