arXiv:2512.00552cs.CL2025-12

小模型看似会解题,实则靠猜——新方法揭穿假推理

Catch Me If You Can: How Smaller Reasoning Models Pretend to Reason with Mathematical Fidelity

  • 设计四维诊断框架,区分真推理与表面匹配
  • 600M模型准确率70%+,但反向一致性仅15%
  • 适合关注模型真实推理能力的研究者

当前语言模型数学推理评估主要依赖答案正确率,可能掩盖逻辑计算的根本缺陷。我们提出一个诊断框架,通过四个互补维度——正向-反向一致性、传递性覆盖度、反事实敏感性、扰动鲁棒性——区分真正的数学推理与表面模式匹配。以Qwen3-0.6B在MenatQA数据集上的案例研究显示,尽管模型答案准确率超过70%,但反向一致性仅为15%,传递性覆盖度仅32.2%,对扰动极为敏感。诊断揭示了传统准确率指标无法捕捉的推理失败,表明该小模型严重依赖模式匹配而非真实逻辑计算。尽管实证基于单一600M参数模型,但该框架具备模型无关性与可扩展性。我们开源评估协议,推动研究界超越表面准确率,实现可验证的数学推理评估。

原文摘要 · Abstract (English)

Current evaluation of mathematical reasoning in language models relies primarily on answer accuracy, potentially masking fundamental failures in logical computation. We introduce a diagnostic framework that distinguishes genuine mathematical reasoning from superficial pattern matching through four complementary axes: forward-backward consistency, transitivity coverage, counterfactual sensitivity, and perturbation robustness. Through a case study applying this framework to Qwen3-0.6B on the MenatQA dataset, we reveal a striking disconnect between surface performance and reasoning fidelity. While the model achieves reasonable answer accuracy (70%+), it demonstrates poor backward consistency (15%), limited transitivity coverage (32.2%), and brittle sensitivity to perturbations. Our diagnostics expose reasoning failures invisible to traditional accuracy metrics, suggesting that this small model relies heavily on pattern matching rather than genuine logical computation. While our empirical findings are based on a single 600M-parameter model, the diagnostic framework itself is model-agnostic and generalizable. We release our evaluation protocols to enable the research community to assess reasoning fidelity across different model scales and architectures, moving beyond surface-level accuracy toward verifiable mathematical reasoning.

数学推理模型诊断小模型真实性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。