不依赖参考答案,检测大模型数学推理中的隐性错误。
A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models
- 通过多维度评估推理步骤可信度、答案一致性与扰动稳定性。
- 在GSM8K和MATH数据集上验证了对错误推理的识别能力。
- 适合用于模型审计、安全评测及训练过程中的隐蔽错误排查。
数学思维链(CoT)评估通常仅以最终答案是否匹配参考答案为标准,这混淆了正确结论与有效推导——无效推理可能偶然得出正确答案,而有效计算也可能因转录错误导致失败。我们称之为‘推理-答案一致性缺口’。本文提出无参考的实例级诊断指标:推理-答案忠实度评分(RAFS),用于评估数学推理轨迹在局部层面是否可信、支持答案,并在重采样与定向反事实干预下保持稳定。RAFS融合步骤有效性、推理到答案的蕴含关系、反事实敏感性、答案共识及条件推理稳定性。它评估的是输出轨迹的一致性,而非模型内部计算或外部事实正确性。研究在GSM8K和MATH上开展预注册、结果盲确认研究,所有假设、准入规则、校准与测试均在观察结果前冻结。另设可行性先导实验,验证端到端执行并估计干预覆盖范围。文中形式化四类推理-答案结果,论证非补偿聚合器合理性,定义语义轨迹距离,量化计算与弃权权衡,提出验证器独立性与功效分析。RAFS旨在补充数学答案准确率,提供可审计的沉默推理失败与答案提取错误预警信号。
原文摘要 · Abstract (English)
Mathematical chain of thought (CoT) evaluation is commonly reduced to whether the final answer matches a reference. This conflates producing a correct conclusion with producing a valid derivation an invalid chain can accidentally reach the right answer, while a valid calculation can be followed by a transcription error. We call this mismatch the reasoning answer consistency gap. This framework paper introduces the Reasoning Answer Faithfulness Score (RAFS), a reference free, instance level diagnostic of whether an emitted mathematical trace is locally credible, supports its answer, and is stable under resampling and targeted counterfactual interventions. RAFS combines step validity, reasoning to answer entailment and counterfactual sensitivity, answer consensus, and conditional reasoning stability. It evaluates transcript level agreement, not a models private computation and not factual correctness outside the tested mathematical setting. We retain a preregistered, results blind confirmatory study on GSM8K and MATH, with hypotheses, admissibility rules, calibration, and tests fixed before confirmatory outcomes are inspected. A separate feasibility pilot is specified to verify end to end execution and estimate interven tion coverage before that freeze numerical pilot claims are re ported only when trace level artifacts are available. We formalize four reasoning answer outcomes, justify the non compensatory aggregator, instantiate semantic trace distance, quantify compute and abstention tradeoffs, and define verifier independence and power analyses. RAFS is intended to complement mathematical answer accuracy with an auditable warning signal for silent reasoning failures and answer extraction errors
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。