用频率响应分析大模型数学推理稳定性,发现精度无法捕捉的动态缺陷。
MathBode: Measuring the Stability of LLM Reasoning using Frequency Response
- 将数学问题建模为系统,通过正弦参数扰动测输出响应
- 揭示模型存在低通滤波特性和相位滞后,精度指标未暴露
- 适合评估模型推理一致性,尤其对高阶应用有指导意义
本文提出 MathBode,一种用于大语言模型(LLM)数学推理的动态诊断方法。不同于一次性准确率评估,MathBode 将每个参数化问题视为一个系统:对单一参数施加正弦扰动,并拟合模型输出与精确解的第一谐波响应。由此获得可解释的、频率解析的度量——增益(幅值跟踪)和相位(延迟),形成类似 Bode 图的指纹。在五类闭式问题(线性求解、比例/饱和、复利、2×2 线性系统、相似三角形)中,诊断揭示出系统性的低通行为和随频率增长的相位滞后,这些特征被传统准确率所掩盖。我们以符号基准模型(G ≈ 1, ϕ ≈ 0)校准工具,对比多个模型,结果在动态特性上清晰区分了前沿与中等水平模型。该方法提供了一种紧凑、可复现的协议,补充标准基准,实现对推理保真度与一致性的可行动度量。数据集与代码已开源,促进后续研究与应用。
原文摘要 · Abstract (English)
This paper presents MathBode, a dynamic diagnostic for mathematical reasoning in large language models (LLMs). Instead of one-shot accuracy, MathBode treats each parametric problem as a system: we drive a single parameter sinusoidally and fit first-harmonic responses of model outputs and exact solutions. This yields interpretable, frequency-resolved metrics -- gain (amplitude tracking) and phase (lag) -- that form Bode-style fingerprints. Across five closed-form families (linear solve, ratio/saturation, compound interest, 2x2 linear systems, similar triangles), the diagnostic surfaces systematic low-pass behavior and growing phase lag that accuracy alone obscures. We compare several models against a symbolic baseline that calibrates the instrument ($G \approx 1$, $ϕ\approx 0$). Results separate frontier from mid-tier models on dynamics, providing a compact, reproducible protocol that complements standard benchmarks with actionable measurements of reasoning fidelity and consistency. We open-source the dataset and code to enable further research and adoption.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。