用微分方程建模推理过程,自动评估大模型推理质量
Markovian ODE-guided scoring can assess the quality of offline reasoning traces in language models
- 基于马尔可夫链与常微分方程建模推理轨迹演化
- 在人类判断下相关性提升超250%(Somers' D)
- 适合评估数学推理、事实核查等高要求任务
生成式语言模型产生的推理轨迹被广泛应用于数学求解、自动化事实核查等任务。然而,现有评估方法仍以机械方式为主,难以捕捉跨任务、渐进退化的推理质量的人类感知维度。本文提出MarODE,一种离线评估框架,可为推理轨迹赋予权重分数。该方法基于推理进展的马尔可夫建模与轨迹动态的常微分方程刻画,实现高效评估。通过人类主导的扰动实验与人工评分,综合验证了评估指标的合理性与有效性。大规模测试显示,MarODE在Somers' D相关性上优于现有基线超过250%。结果表明,理论驱动的评估框架对日益关键的推理轨迹评估具有重要意义。
原文摘要 · Abstract (English)
Reasoning traces produced by generative language models are increasingly used for tasks ranging from mathematical problem solving to automated fact checking. However, existing evaluation methods remain largely mechanical and fail to capture human-centric notions of reasoning quality in a way that generalizes across varied and progressively degraded reasoning. We introduce MarODE, an offline evaluation framework that assigns quality scores to reasoning traces. Its effectiveness is assessed using human-centric perturbations and human judgments, which jointly evaluate the fundamental dimensions of an evaluation metric - goodness and soundness. The approach is grounded in a Markovian formulation of reasoning progression and an ordinary differential equation based characterization of trace dynamics, enabling efficient evaluation of reasoning quality. In a large-scale evaluation, MarODE outperforms existing baselines by over 250% under Somers' D correlation. Our results emphasize the value of theory-driven evaluation frameworks as reasoning traces become central to language model-based systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。