重新审视语言模型的完成态误判问题,发现是评测设计缺陷而非模型固有偏见。
The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure
- 通过构建词义匹配的最小对比对,修正评测中的概念混淆问题。
- 发现模型在未明确完成时仍接受完成态结论,体现为充分性偏差。
- 提示干预导致判断转移但未提升真实理解,适合语言学与评测研究者参考。
imperfective paradox 为组合语义分析提供有效测试。近期研究构建了一个 NLI 基准并报告模型常从进行时描述推断完成的有界事件,归因于目的性偏见,并认为提示干预引发校准危机。本文重新审视该基准及结论,指出其受概念与评估设定错误的显著影响。我们识别出三项概念误设,特别是体貌缩减(Aspectual Reduction)影响了基准构建、分析、实验与结论。在严格 NLI 标准下,76% 的 Group A 实例未明确排除结果达成。母语者标注显示,38% 的 Group A 和 29% 的 Group C 示例允许多种解释。为控制此问题与词汇变异,我们构建了词义匹配的最小对比对。在评估层面,将事件语义 NLI 视为多步推理问题,评估中间语义判断与最终预测。结果表明,模型常不确认结果达成,却仍接受简单过去时假设,这一现象称为充分性偏差。提示干预引发标签间的决策转移,但未可靠提升底层语义理解与推理能力。中间与最优引导分析揭示两种额外失败模式:组合体貌分类错误,以及表面形式吸引至表面关联答案。对 Qwen-7B、GPT-5.4 与 Qwen-72B 的实验提供初步证据,表明体貌分类具有上下文敏感性,且这些模型在适当提示下可达到与人类标注者相当的表现。
原文摘要 · Abstract (English)
The imperfective paradox provides a useful test of compositional semantic analysis. Recent work constructs an NLI benchmark and reports that models frequently infer completed telic events from progressive descriptions, attributing this behavior to a Teleological Bias. It further argues that prompting interventions cause a Calibration Crisis. We reexamine the benchmark and conclusions and show that it is substantially affected by conceptual and evaluation mis-specifications. We identify three conceptual mis-specifications. In particular, Aspectual Reduction affects the benchmark construction, analysis, experiments, and conclusions. Under a strict NLI standard, 76% of Group A instances do not explicitly rule out culmination. In our native-speaker annotation, 38% of Group A examples and 29% of the Group C examples were judged to permit an alternative interpretation. To control these issues and lexical variation, we construct Lexically Matched Minimal Pairs. At the evaluation level, we formulate event-semantic NLI as a Multi-step Reasoning Problem and assess both intermediate semantic decisions and final predictions. Our results show that models often do not affirm culmination but nevertheless accept the corresponding simple-past hypothesis, a pattern we characterize as Sufficiency Bias. We further show that prompting interventions produce a Decision Shift among labels without reliably improving the underlying semantic understanding and reasoning. Intermediate and oracle-guided analyses identify two additional failure modes: errors in compositional aspectual classification and Surface-form Attraction toward surface-associated answers. Our experiments on Qwen-7B with suitable prompts, GPT-5.4, and Qwen-72B provide initial evidence for the context sensitivity of aspectual classification and suggest that these models can achieve performance comparable to that of human annotators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。