arXiv:2603.05290cs.AI2026-03KDD

用形式化探针精准测量大模型推理能力,发现其对约束变化敏感但结构重构易失效。

X-RAY: Mapping LLM Reasoning Capability via Formalized and Calibrated Probes

  • 通过形式化工具生成可校准的探针,分离结构化推理信息
  • 模型在解空间重构时性能骤降,而约束细化下表现稳定
  • 可识别传统评测中难以发现的结构性失败模式,适合研究者使用

大语言模型虽表现优异,但其推理能力仍不清晰。现有评估多关注任务准确率,常将模式匹配与真实推理混淆。本文提出X-RAY,一种可解释的推理分析系统,通过形式化且可校准的探针映射大模型的推理能力。我们将推理能力建模为可提取结构的函数,通过约束交互、推理深度和解空间几何等正式属性进行操作化。X-RAY利用形式工具生成具有可控结构变化的探针,实现增量结构信息的精确隔离。我们在数学、物理和化学领域从初级到高级的问题上评估了先进大模型。分析揭示出模型推理的系统性不对称:对约束细化(缩小解空间)相对稳健,但在解空间重构(改变解流形结构)时性能急剧下降。此外,校准的形式探针能区分传统基准上表现相似的模型,并揭示结构可解释而非模糊的失败模式。该框架还支持无污染的训练与测试,适用于推理模型开发。

原文摘要 · Abstract (English)

Large language models (LLMs) achieve promising performance, yet their ability to reason remains poorly understood. Existing evaluations largely emphasize task-level accuracy, often conflating pattern matching with reasoning capability. We present X-RAY, an explainable reasoning analysis system that maps the LLM reasoning capability using calibrated, formally verified probes. We model reasoning capability as a function of extractable \textit{structure}, operationalized through formal properties such as constraint interaction, reasoning depth, and solution-space geometry. X-Ray generates probes via formal tools with controlled structural variations, enabling precise isolation of incremental structural information through formal calibration and verification. We evaluate state-of-the-art LLMs on problems ranging from junior-level to advanced in mathematics, physics, and chemistry. Our analysis reveals a systematic asymmetry in LLM reasoning: models are relatively robust to constraint refinement, where additional conditions shrink an existing solution space, but degrade sharply under solution-space restructuring, where modifications alter the underlying structural form of the solution manifold. Moreover, calibrated formal probes differentiate models that appear indistinguishable on standard benchmarks and reveal failure modes that are structurally interpretable rather than opaque. Beyond evaluation, our framework is contamination-free and supports the training and testing of reasoning models.

大模型推理形式化探针结构分析可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。