评测大模型对动态系统逻辑推理的短板,发现多数模型表现接近随机。
ChaosBench-Logic v2: Evaluating LLM Logical Reasoning over Dynamical Systems at Scale

- 构建4万道题的动态系统逻辑推理基准,含27个一阶逻辑谓词
- 前沿模型在参数依赖推理上准确率仅0.05(接近随机)
- 开源模型在指标诊断任务中领先,但部分模型出现系统性反相关
标准二分类评测掩盖了关键失败模式:先前崩溃、改写不一致以及对参数依赖动力学的推理失效。我们提出ChaosBench-Logic v2,一个包含40,886道题、165个动力系统、27个一阶逻辑(FOL)谓词和78条公理边的大规模基准,配套开发了CARE(校准与对抗鲁棒评估)协议以揭示这些病理。评估14个模型发现,即使在最先进模型中,相变推理仍接近随机(MCC = 0.05),而给定前提的一阶逻辑推导可达MCC = 0.52。按系统家族分解显示,专有模型优势集中于跨指标(+0.40)与一致性任务,而开源Qwen 2.5-32B在指标诊断中表现突出(0.91对比0.45)。两个模型在分岔问题上呈现负MCC,混淆矩阵分析确认其存在系统性反相关。
原文摘要 · Abstract (English)
Standard accuracy on binary reasoning benchmarks hides critical failure modes: prior collapse, inconsistency under paraphrase, and inability to reason about parameter-dependent dynamics. We present ChaosBench-Logic v2, a 40,886-question benchmark over 165 dynamical systems with 27 FOL predicates and 78 axiom edges, together with CARE (Calibration- and Adversarial-Robust Evaluation), a protocol that surfaces these pathologies. Evaluating 14 models, we find that regime-transition reasoning remains near random (MCC = 0.05) even for frontier models, whereas FOL deduction with given premises reaches MCC = 0.52. Per-family decomposition shows that the proprietary-model advantage concentrates on cross-indicator (+0.40) and consistency tasks, while open-source Qwen 2.5-32B dominates indicator diagnostics (0.91 vs. 0.45). Two models exhibit negative MCC on bifurcation questions, confirmed as systematic anti-correlation via confusion-matrix analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。