arXiv:2608.30256cs.CLcs.AI2026-08中稿 · EMNLP

用符号编辑测试大模型逻辑推理可靠性,发现其常忽略结构变化。

Beyond Surface Forms: Symbolic Edits as a Test for Logical Reasoning with LLMs

论文配图:Beyond Surface Forms: Symbolic Edits as a Test for Logical Reasoning with LLMs
图 1 · 摘自论文原文
  • 将逻辑题转为符号形式,精准修改运算符后还原成自然语言
  • 模型对逻辑结构改变反应不一,大模型也常犯错
  • 适合评估大模型真实推理能力,尤其关注逻辑一致性

大型语言模型(LLMs)的逻辑推理能力至关重要,体现其能否基于上下文正确推导假设并遵循忠实的演绎过程。然而,现有研究显示,模型对问题表述中的细微表面变化极为敏感,引发对其是否真正理解底层逻辑结构的质疑。由于自然语言中逻辑成分(如算子、谓词)难以系统操控,该行为研究极具挑战。本文提出一种工具驱动的框架,可对一阶逻辑和约束满足问题任务的符号表示进行可控、标签保持的编辑,实现对逻辑算子等结构组件的针对性修改,并在转换回自然语言前完成重构。利用此框架,我们在累积与单个算子编辑下评估多个LLMs的行为,并分析其响应。定量与定性分析表明,无论模型规模或家族如何,其推理行为在受控算子编辑下均表现出不一致性:模型有时能正确适应结构变化,但更多时候无法追踪其逻辑后果。该自动化压力测试结果可多维度评估语言模型,有助于衡量其推理可靠性。

原文摘要 · Abstract (English)

Logical reasoning with large language models (LLMs) is a critical capability, as it reflects a system's ability to correctly deduce hypotheses from a given context using faithful deductive processes. However, LLM reasoning has often been shown to be sensitive to small surface-level variations in problem formulation, raising questions about whether models truly follow the underlying logical structure. Studying this behavior is challenging because the symbolic components of logical problems, such as operators and predicates, are difficult to systematically manipulate in natural language. We introduce a tool-driven framework for generating controlled, label-preserving edits to logical reasoning problems. Our method operates on symbolic representations of first-order logic and constraint satisfaction problem tasks, enabling targeted modifications to logical operators and other structural components before translating them back into natural language. Using this framework, we evaluate various LLMs under cumulative and individual operator edits and analyze their behavior in response to these changes. Our quantitative and qualitative analyses show that LLM reasoning behavior under controlled operator edits is inconsistent, regardless of model size or family: models sometimes adapt correctly to structural changes but often fail to track their logical consequences. The results from this automated stress test enable an evaluation of language models across different dimensions and help measure the reliability of their reasoning.

逻辑推理大模型评估符号编辑可靠性测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。