arXiv:2608.05097cs.CL2026-08

测试大模型是否真懂模态逻辑,发现推理模式决定表现

Same Formulas, Different Semantics: Do Language Models Follow Modal Logic Specifications?

  • 设计成对的模态题,仅语义条件不同但形式相同
  • 五模型中四款在直接提示下准确率低于基线,最高仅4.4%
  • 开启推理模式后,深思V4闪现提升至88.1%,体现推理模式关键作用

关于必然性与可能性的推理依赖于世界间可及性及各世界中存在的对象。同一推理在一种模态系统中成立,可能在另一系统中不成立。评估语言模型需检验其判断是否遵循指定语义,而非熟悉逻辑。我们构建了前提和结论相同的成对模态问题,但框架或域条件不同;自动推理验证得出相反标签。平衡核心确保仅凭语义条件无法揭示答案。在此核心上,五个近期模型中有四个在直接提示下表现低于仅依赖语义条件的基线。然而,启用推理模式后,DeepSeek V4 Flash在相同提示下准确率从4.4%提升至88.1%。遵循指定模态语义强烈依赖于推理模式与模型身份。当框架条件缺失时,模型常达成一致,但更符合不同熟悉逻辑。代码、公式、验证工具、反例与响应均已公开。

原文摘要 · Abstract (English)

Reasoning about necessity and possibility depends on assumptions about accessibility between worlds and about which objects exist at each one. The same inference may therefore hold under one modal system and fail under another. Evaluating language models on such problems requires testing whether their judgments follow the stated semantics rather than a familiar logic. We construct paired modal problems with identical premises and conjecture but different frame or domain conditions; automated reasoning verifies opposite labels. A balanced core prevents the semantic condition alone from revealing the answer. On this core, four of five recent models perform below the condition-only baseline under direct prompting. Yet enabling reasoning mode raises DeepSeek V4 Flash from 4.4% to 88.1% on unchanged prompts. Following stipulated modal semantics thus depends strongly on inference mode as well as model identity. When frame conditions are omitted, models often agree but fit different familiar logics best. We release the formulas, oracle artifacts, countermodels, and responses.

模态逻辑大模型推理语义理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。