arXiv:2509.22518cs.AIcs.LG2025-09被引 3

通过几何结构分析,定位大模型推理失败的根源。

REMA: A Unified Reasoning Manifold Framework for Interpreting Large Language Model

  • 构建推理流形,用低维几何结构表征正确推理路径。
  • 量化错误样本与正确流形的距离,识别失败发生位置。
  • 适用于诊断大模型推理错误,适合可解释性研究者。

理解大语言模型(LLMs)如何执行复杂推理及其失效机制是可解释性研究中的挑战。为此,我们提出‘推理流形’概念——由所有正确推理生成对应的内部表示构成的低维几何结构,可视为模型学会的有效思维路径的体现。基于此,我们构建了REMA框架,通过定量比较错误与正确推理样本的内部表示空间关系,解释失效原因。REMA首先计算每个错误表示到正确表示近似流形的k近邻距离,获得统一的失败信号;再沿模型各层追踪该偏差,并与正确表示的内部波动基线对比,定位推理链偏离的初始节点。在多种语言和多模态模型及任务上的广泛实验验证了推理流形的低维特性,以及错误与正确表示的高度可分性。结果证明了REMA在分析推理失败起源方面的有效性。该研究将抽象推理失败映射到表示空间的可测量几何偏移,为深入理解黑箱模型内部计算过程提供了新途径。

原文摘要 · Abstract (English)

Understanding how Large Language Models (LLMs) perform complex reasoning and their failure mechanisms is a challenge in interpretability research. To provide a measurable geometric analysis perspective, we define the concept of the Reasoning Manifold, a latent low-dimensional geometric structure formed by the internal representations corresponding to all correctly reasoned generations. This structure can be conceptualized as the embodiment of the effective thinking paths that the model has learned to successfully solve a given task. Based on this concept, we build REMA, a framework that explains the origins of failures by quantitatively comparing the spatial relationships of internal model representations corresponding to both erroneous and correct reasoning samples. Specifically, REMA first quantifies the geometric deviation of each erroneous representation by calculating its k-nearest neighbors distance to the approximated manifold formed by correct representations, thereby providing a unified failure signal. It then localizes the divergence points where these deviations first become significant by tracking this deviation metric across the model's layers and comparing it against a baseline of internal fluctuations from correct representations, thus identifying where the reasoning chain begins to go off-track. Our extensive experiments on diverse language and multimodal models and tasks demonstrate the low-dimensional nature of the reasoning manifold and the high separability between erroneous and correct reasoning representations. The results also validate the effectiveness of the REMA framework in analyzing the origins of reasoning failures. This research connects abstract reasoning failures to measurable geometric deviations in representations, providing new avenues for in-depth understanding and diagnosis of the internal computational processes of black-box models.

模型可解释性推理分析几何方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。