arXiv:2606.19636cs.LGcs.AI2026-06被引 1

发现数学推理难题的采样盲区,提出新诊断方法揭示模型隐藏能力。

Hard or Just Unreached? Diagnosing the Sampling Blind Spot in Math-Reasoning Difficulty Estimation

  • 用激活嫁接技术扰动模型内部表示,提升解题多样性。
  • 在六次采样均失败的难题中,10.3%-22.9%可被确定性解法攻克。
  • 适用于评估大模型数学推理真实能力,适合模型优化与数据筛选。

数学与科学推理基准依赖pass@k(采样链中达到正确答案的比例)作为每题难度的核心信号,该信号驱动强化学习、数据筛选、合成课程与验证器训练。我们发现该代理指标在最困难题型上存在持续盲区:在八组自由格式数学题(GSM8K与MATH,四款开源模型)中,6次采样均未解出的题目中,有10.3%-22.9%可通过六条链的确定性策略成功解决。该策略为贪婪解码加五次廉价残差流扰动,通过激活嫁接实现;而仅贪婪解码在这些题目上最多解决6%。恢复能力随额外计算预算增加而提升,且不同扰动方式的机制差异经验证(跨类别修复集杰卡德系数≤0.47)。激活嫁接仅用于干预内部表示,非解码方法;我们仅将其作为诊断与多样化工具,证明pass@k=0%的题目在残差流中具有结构性可识别性,而非原模型在常规推理下无法触及。

原文摘要 · Abstract (English)

Math and science reasoning benchmarks rely on pass@k, the fraction of sampled chains that reach gold, as the canonical per-example difficulty signal. The same signal drives RL with verifiable rewards, math data curation, synthetic curricula, and verifier training. We show this proxy has a persistent blind spot on its hardest stratum: on the eight free-form math cells we test (GSM8K and MATH across four open-weight models), 10.3-22.9% of the examples that no sampling seed solves in six tries are instead solved at matched compute by a six-chain deterministic regime. These are greedy decoding plus five cheap residual-stream perturbations applied via activation grafting, while greedy alone solves at most 6% on these math cells. Recovery scales with the additional budget, across perturbations whose mechanistic distinctness we verify across all twelve cells (cross-kind fix-set Jaccard <= 0.47 in every setup). Activation grafting is used as an intervention on internal representations, not a decoding method; we use it purely as a diagnostic and diversification tool, and our recovered items show that the pass@k= 0 % stratum is structurally identifiable in the residual stream rather than that the unmodified model reaches them under ordinary inference.

数学推理模型诊断采样盲区激活嫁接

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。