arXiv:2505.13408cs.AIcs.CL2025-05被引 62

用物理动力学建模大模型推理过程,量化推理可信度。

CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process

  • 将推理过程类比为粒子在力学场中的运动,构建可评分的能量方程。
  • 首次实现对推理路径合理性进行独立评分,提升答案可信度判断精度。
  • 适合关注模型可解释性与推理质量评估的研究者。

近期大型推理模型(LRM)通过学习推理能力,显著提升了复杂任务求解性能。这些模型通过生成显式的推理轨迹来回答问题,但仅判断答案正确性不足以评估输出质量,推理过程的逻辑严谨性同样关键。若推理过程不严谨,即使答案正确,其置信度也应较低。现有方法虽尝试联合评估推理与答案,但难以准确反映推理对结论的因果影响。本文受经典力学启发,提出一种新型的 CoT-Kinetics 能量方程,将大模型内部 Transformer 层驱动的标记状态演化建模为受力场支配的动力学过程。该能量方程为推理阶段分配标量得分,精确衡量推理的合理性,从而更准确地评估模型整体输出质量,突破传统‘正确/错误’的粗粒度判断局限。

原文摘要 · Abstract (English)

Recent Large Reasoning Models significantly improve the reasoning ability of Large Language Models by learning to reason, exhibiting the promising performance in solving complex tasks. LRMs solve tasks that require complex reasoning by explicitly generating reasoning trajectories together with answers. Nevertheless, judging the quality of such an output answer is not easy because only considering the correctness of the answer is not enough and the soundness of the reasoning trajectory part matters as well. Logically, if the soundness of the reasoning part is poor, even if the answer is correct, the confidence of the derived answer should be low. Existing methods did consider jointly assessing the overall output answer by taking into account the reasoning part, however, their capability is still not satisfactory as the causal relationship of the reasoning to the concluded answer cannot properly reflected. In this paper, inspired by classical mechanics, we present a novel approach towards establishing a CoT-Kinetics energy equation. Specifically, our CoT-Kinetics energy equation formulates the token state transformation process, which is regulated by LRM internal transformer layers, as like a particle kinetics dynamics governed in a mechanical field. Our CoT-Kinetics energy assigns a scalar score to evaluate specifically the soundness of the reasoning phase, telling how confident the derived answer could be given the evaluated reasoning. As such, the LRM's overall output quality can be accurately measured, rather than a coarse judgment (e.g., correct or incorrect) anymore.

推理评估模型可信度动力学建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。