arXiv:2509.10860cs.CL2025-09被引 1

对比中英文下大模型与人类对量词作用域的理解差异

Quantifier Scope Interpretation in Language Learners and LLMs

  • 用概率评估大模型在中英文中的量词作用域解读偏好
  • 多数模型倾向表层作用域,部分能区分英中文逆序偏好
  • 预训练数据语言背景显著影响模型类人表现

含多个量词的句子常引发解释歧义,且跨语言存在差异。本研究采用跨语言方法,考察大语言模型(LLMs)在英语和汉语中处理量词作用域解读的能力,使用概率评估解释可能性。通过人类相似性(HS)分数量化模型模拟人类表现的程度。结果显示,大多数LLMs偏好表层作用域解读,符合人类倾向;仅部分模型能在英中文之间区分逆序作用域偏好,反映类人模式。HS分数揭示了模型类人行为的差异性,但整体上表现出与人类对齐的潜力。模型架构、规模及特别是预训练数据的语言背景,显著影响其接近人类量词作用域解读的程度。

原文摘要 · Abstract (English)

Sentences with multiple quantifiers often lead to interpretive ambiguities, which can vary across languages. This study adopts a cross-linguistic approach to examine how large language models (LLMs) handle quantifier scope interpretation in English and Chinese, using probabilities to assess interpretive likelihood. Human similarity (HS) scores were used to quantify the extent to which LLMs emulate human performance across language groups. Results reveal that most LLMs prefer the surface scope interpretations, aligning with human tendencies, while only some differentiate between English and Chinese in the inverse scope preferences, reflecting human-similar patterns. HS scores highlight variability in LLMs' approximation of human behavior, but their overall potential to align with humans is notable. Differences in model architecture, scale, and particularly models' pre-training data language background, significantly influence how closely LLMs approximate human quantifier scope interpretations.

语言理解大模型量词歧义

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。