测试大模型在逻辑等价改写下的答案一致性,发现高准确率不等于逻辑自洽。
Controlled Reformulation Testing for Logical Consistency in Large Language Models

- 设计350组逻辑等价问题,涵盖逆否、双重否定等7类改写方式。
- GPT-5.4-mini准确率达98.9%但一致性仅60.3%,推理优化模型达96.9%。
- 复杂逻辑变换如逆否命题失败率超70%,表面重述仍稳定可靠。
大语言模型在逻辑等价问题形式变化时经常自相矛盾。本文提出包含350个问题族(共1,750个问题)的可控改写测试基准(CRTBench),评估模型在逆否命题、双重否定、否定翻转和被动语态等控制改写下的答案一致性。实验发现,尽管GPT-5.4-mini基线准确率达98.9%,但家族级一致性仅为60.3%;而推理优化的o4-mini达到96.9%的一致性。失败主要集中在非平凡逻辑变换:逆否命题失败率72.4%,双重否定为84.6%;表面重述则保持94–100%稳定。增加推理努力使GPT-5.4-mini一致性提升至85.4%,但整体未变,因嵌套否定增益被量词类失败抵消。结果表明,仅靠准确率无法评估模型逻辑推理能力。
原文摘要 · Abstract (English)
Large language models (LLMs) frequently contradict themselves when the surface form of a logically equivalent question changes. We present a benchmark of 350 question families (1,750 total questions) for Controlled Reformulation Testing (CRTBench) to evaluate logical invariance. In this benchmark, we investigate LLMs' ability to maintain consistent answers across controlled reformulations, which include contrapositive rewriting, double negation, negation flipping, and passive voice. We evaluate several frontier LLMs and observe an accuracy-consistency gap where GPT-5.4-mini achieves $98.9\%$ base accuracy but only $60.3\%$ family-level consistency, while reasoning-optimized o4-mini achieves $96.9\%$ consistency. From our experiments, we observe that failures cluster around logically nontrivial transformations such as contrapositive rewriting ($72.4\%$ for GPT-5.4-mini) and double negation ($84.6\%$), while surface-level rephrasing remains robust ($94-100\%$). Increasing reasoning effort improves GPT-5.4-mini to $85.4\%$ consistency, but leaves GPT-5.4 unchanged overall because gains on nested negation are offset by failures on quantifier families. These results show that accuracy alone is not enough for evaluating logical reasoning in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。