对比人类与大模型在条件句预设投射上的判断差异
Presupposition and Reasoning in Conditionals: A Theory-Based Study of Humans and LLMs

- 设计对照实验,比较人类与4个大模型对条件句预设的判断
- 人类结合概率和语用线索,大模型表现不一且多依赖表面模式
- 发现推理能力强的模型反而更不像人,提示当前评估需理论支撑
条件句中的预设投射是语义与语用理论的核心问题,但大型语言模型在此方面的表现尚未得到充分评估。我们通过一项平行行为研究,比较人类判断与大模型在一组经过规范化的条件句上的预测结果,该数据集控制了前件与预设投射之间的关系。我们收集了120名参与者和四个大模型在相同上下文条件下的似然度评分。结果显示,人类在判断中融合了概率与语用线索,而大模型的表现则与人类模式存在差异。通过在大模型作为裁判的框架内使用语言学驱动的检查清单进一步评估模型推理,我们发现那些最接近人类评分的模型往往缺乏连贯的语用推理能力,而推理能力较强的模型则产生更不类人的判断。这些发现表明,大模型在该任务上的表现可能源于表面模式匹配而非真正的语用理解。研究强调了基于语言学理论的基准测试对于人类与模型对比的重要性。
原文摘要 · Abstract (English)
Presupposition projection in conditionals is central to theories of meaning and pragmatics, yet it remains largely unevaluated in large language models. We address this gap through a parallel behavioral study comparing human judgments and LLM predictions on a normed dataset of conditional sentences that controls the relation between the antecedent and the projected presupposition. We collect likelihood ratings from 120 participants and four LLMs under matched contextual conditions. Results show that humans integrate probabilistic and pragmatic cues in their judgment, whereas LLMs show variable alignment with human patterns. Using a linguistically motivated checklist within an LLM-as-a-Judge framework, we further evaluate model reasoning. We observe models that best match human ratings often lack coherent pragmatic reasoning, while models with stronger reasoning produce less human-like judgments. These findings suggest that LLMs' performance on such tasks may result from surface pattern matching rather than pragmatic competence. Our findings highlight the importance of benchmarks grounded in linguistic theory for comparing humans and models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。