用逻辑组合强化测试,暴露大模型在复杂推理中的系统性缺陷。
From 0-Order Selection to 2-Order Judgment: Combinatorial Hardening Exposes Compositional Failures in Frontier LLMs
- 将选择题转化为多步逻辑判断题,提升推理难度与思考深度。
- 12个顶尖模型准确率下降31%至56%,零样本迁移降幅达47%。
- 适合研究模型推理瓶颈、评估真实思维能力的研究者使用。
多选题推理基准面临两大挑战:模型进步导致成绩快速饱和,以及数据污染削弱静态评测有效性。现有临时增强方法(如改写、扰动)虽增加难度,却以牺牲逻辑一致性为代价,难以真正挑战先进推理模型。本文提出LogiHard框架,可确定性地将0阶选择题转化为2阶逻辑判断题,显著增加思维开销与推理步骤。该框架结合项目反应理论(IRT)实现自适应测试(CAT),用更少题目精准控制难度。我们构建了LogiHard-2k数据集,通过9维分析模型思维轨迹,对高难度考试题进行认知排序,并对高难题目实施组合变换。在12个前沿模型上评估显示,组合强化题目下准确率下降31%至56%。大模型存在多选失败和过早退出偏差,而人类考生无此问题。零样本迁移至MMLU测试中准确率从89.84%降至42.86%(下降47%),证实跨领域适用性与有效性保持。整体性能下降具有领域无关性,非知识不足所致,而是源于组合推理能力缺失,反映训练中完整性验证机制的缺陷。
原文摘要 · Abstract (English)
Multiple-choice reasoning benchmarks face dual challenges: rapid saturation from advancing models and data contamination that undermines static evaluations. Ad-hoc hardening methods (paraphrasing, perturbation) attempt to increase difficulty but sacrifice logical validity for surface complexity, falling short to challenge advanced reasoning models. We present LogiHard, a formal framework that deterministically transforms 0-order selection into 2-order logical judgment, which significantly increases the thinking overhead and reasoning steps. The framework integrates Item Response Theory (IRT) for computerized adaptive testing (CAT), enabling precise difficulty control with fewer questions than static benchmarks. We instantiate LogiHard-2k, a logical reasoning dataset constructed by cognitively ranking high-stakes examination questions via 9-dimensional analysis of model thinking traces, followed by combinatorial transformation of high-difficulty items. Evaluation across twelve state-of-the-art models reveals an accuracy degradation ranging from 31% to 56% on combinatorially hardened questions. LLMs suffer from the multi-select failure and early exit bias, which are not shared by human testees. Zero-shot transfer to MMLU demonstrates 47% accuracy degradation (89.84% to 42.86%), confirming applicability across domains with provable validity preservation. The consistent aggregate degeneration is domain-agnostic and stems not from knowledge deficits but from a combinatorial reasoning gap, reflecting a training-induced completeness-verification deficit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。