arXiv:2602.12665cs.AI2026-02被引 1

用可调控逻辑结构测试大模型推理鲁棒性,发现表面不变时仍会突然失效。

Evaluating Robustness of Reasoning Models on Parameterized Logical Problems

  • 设计参数化2-SAT基准,通过图结构控制可满足性
  • 模型在结构扰动下性能骤降,即使表面特征不变
  • 适合评估大模型逻辑推理的稳定性与抗干扰能力

逻辑为评估基于大模型的推理器提供了受控环境,但传统SAT类基准常将表面难度(长度、措辞、子句顺序)与决定可满足性的结构性因素混淆。我们提出一种基于参数化结构化2-CNF公式的诊断性2-SAT基准,其可满足性由蕴含图决定,可通过可解释轴进行调节。生成器分离出不同能力与失败模式:(i) 可控大小和失衡的矛盾环不可满足核心;(ii) 指定自由变量比例以控制解的多重性;(iii) 植入骨架以调节传播;(iv) 延迟桥接子句耦合原本单调区域,探测对顺序和修正的敏感性;(v) 对称/重复变体测试重命名与冗余结构下的抽象能力。我们在决策准确性和赋值有效性上评估基于大模型的推理器,并量化其在语义保持扰动(如子句重排、填充子句、变量重命名)下的鲁棒性。在多个模型中,我们观察到在特定结构干预下性能出现急剧转变,即使表面统计量固定,揭示了聚合SAT准确率无法察觉的脆弱区间。

原文摘要 · Abstract (English)

Logic provides a controlled testbed for evaluating LLM-based reasoners, yet standard SAT-style benchmarks often conflate surface difficulty (length, wording, clause order) with the structural phenomena that actually determine satisfiability. We introduce a diagnostic benchmark for 2-SAT built from parameterized families of structured 2--CNF formulas, where satisfiability is characterized by the implication graph and can be tuned along interpretable axes. Our generators isolate distinct competencies and failure modes: (i) contradiction-cycle UNSAT cores with controllable size and imbalance, (ii) SAT instances with a prescribed fraction of free variables to control solution multiplicity, (iii) planted backbones that modulate propagation, (iv) late bridge clauses that couple otherwise monotone regions to probe sensitivity to ordering and revision, and (v) symmetry/duplication variants that test abstraction under renaming and redundant structure. We evaluate LLM-based reasoners on decision accuracy and assignment validity, and quantify robustness under semantics-preserving perturbations such as clause reordering, filler clauses, and variable renaming. Across models, we observe sharp performance transitions under targeted structural interventions even when surface statistics are held fixed, revealing brittleness regimes that are invisible to aggregate SAT accuracy.

逻辑推理鲁棒性评估2-SAT大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。