arXiv:2601.09373cs.CL2026-01ACL被引 1

LLM在时态推理中常错误认为目标行为已完成,暴露其逻辑推理缺陷。

The Imperfective Paradox in Large Language Models

  • 构建诊断数据集ImperfectiveNLI,测试模型对进行时与完成时的语义理解
  • 发现模型普遍存在目标导向偏见,即使文本明确否定仍误判完成
  • 适合研究模型语义理解、认知偏差或可解释性的研究人员参考

大型语言模型是否真正理解事件的组合语义,还是仅依赖表层概率启发式?我们研究了‘未完成性悖论’——过去进行时对活动类(如跑步→跑完)蕴含事件实现,但对成就类(如建造→建完)却不蕴含。为此,我们提出ImperfectiveNLI诊断数据集,用于探测不同语义类别中的这一区别。评估当前主流开源模型发现,模型普遍存在目的论偏差:对目标导向事件系统性地错判为已完成,即使文本中已明确取消。提示干预虽部分缓解该偏差,却引发校准危机,导致模型错误拒绝非目的性动词的合法蕴含关系。表示分析进一步显示,尽管内部嵌入能区分进行时与简单过去时形式,但推理决策仍受强烈的目标达成先验支配。综合结果表明,当前开源大模型更像预测性叙事引擎而非忠实逻辑推理者,而解决时态推理需超越提示工程,走向结构化对齐。

原文摘要 · Abstract (English)

Do Large Language Models (LLMs) genuinely grasp the compositional semantics of events, or do they rely on surface-level probabilistic heuristics? We investigate the Imperfective Paradox, a logical phenomenon where the past progressive aspect entails event realization for activities (e.g., running $\to$ ran) but not for accomplishments (e.g., building $\nrightarrow$ built). We introduce ImperfectiveNLI, a diagnostic dataset designed to probe this distinction across diverse semantic classes. Evaluating state-of-the-art open-weight models, we uncover a pervasive Teleological Bias: models systematically hallucinate completion for goal-oriented events, even overriding explicit textual cancellation. Prompting interventions partially reduce this bias but trigger a calibration crisis, causing models to incorrectly reject valid entailments for atelic verbs. Representational analyses further show that while internal embeddings often distinguish progressive from simple past forms, inference decisions are dominated by strong priors about goal attainment. Taken together, our findings indicate that these current open-weight LLMs operate as predictive narrative engines rather than faithful logical reasoners, and that resolving aspectual inference requires moving beyond prompting toward structurally grounded alignment.

语言模型语义理解认知偏差逻辑推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。