arXiv:2603.08358cs.CL2026-03被引 1

测试大模型能否理解条件句中的隐含前提,发现它们靠模式匹配而非真正理解。

Do Language Models Know Theo Has a Wife? Investigating the Proviso Problem

  • 将隐含前提问题转为自然语言推理任务,构建诊断数据集。
  • 模型判断与人类一致,但依赖表面模式而非深层语用推理。
  • 首次提供计算评估框架,适合研究语用能力的学者参考。

我们研究语言模型如何处理条件句中隐含前提的投射问题,即理论解释与人类理解之间的分歧。将该现象重构为自然语言推理任务,并设计了一个诊断数据集以探测条件句中的前提投射。使用RoBERTa、DeBERTa、LLaMA和Gemma进行评估,并结合可解释性分析。结果表明,模型整体判断与人类一致,但依赖浅层模式匹配,而非语义或语用推理。本工作首次建立了针对该问题的计算评估框架,强调需采用诊断性、多方法手段来评估语言模型的语用能力和上下文依赖意义理解。

原文摘要 · Abstract (English)

We investigate how language models handle the proviso problem, an unresolved issue in pragmatics where presuppositions in conditional sentences diverge between theoretical and human interpretations. We reformulate this phenomenon as a Natural Language Inference task and introduce a diagnostic dataset designed to probe presupposition projection in conditionals. We evaluate RoBERTa, DeBERTa, LLaMA, and Gemma using explainability analyses. The results show that models broadly align with human judgments but rely on shallow pattern matching rather than semantic or pragmatic reasoning. Our work provides the first computational evaluation framework for the proviso problem and highlights the need for diagnostic, multi-method approaches to assess pragmatic competence and context-dependent meaning in language models.

语用学前提投射大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。