arXiv:2512.19135cs.AI2025-12被引 2

用拓扑分析方法揭示大模型推理链的结构质量与准确性的关系。

Understanding Chain-of-Thought in Large Language Models via Topological Data Analysis

  • 通过拓扑数据分析推理步骤的语义结构,提取连通性与冗余度特征。
  • 复杂度高的推理链更早找到正确答案,成功推理具有更简洁拓扑结构。
  • 为优化推理链效率和可解释性提供新视角,适合关注模型可解释性的研究者。

随着大语言模型(LLMs)的发展,特别是长推理链技术的引入,复杂问题求解中的推理能力显著提升。然而,不同推理链表现差异的原因尚不明确,其关键结构成分亦未被充分理解。现有研究多从功能角度评估推理链,忽视其内在结构机制。本文首次从结构视角分析推理链质量,采用拓扑数据分析(TDA)中的持久同调(persistent homology)将推理步骤映射到语义空间,提取拓扑特征并分析结构变化,揭示语义连贯性、逻辑冗余及断点与空白。通过计算同调群评估多尺度下的连通性与冗余度,并利用条形图与持久图量化稳定性与一致性。结果表明,推理链的拓扑结构复杂度与准确性正相关:复杂链更早定位正确答案,而成功推理表现为更简单的拓扑结构,减少冗余与环路,提升效率与可解释性。本工作为推理链质量评估提供了新视角,并为未来优化提供指导。

原文摘要 · Abstract (English)

With the development of large language models (LLMs), particularly with the introduction of the long reasoning chain technique, the reasoning ability of LLMs in complex problem-solving has been significantly enhanced. While acknowledging the power of long reasoning chains, we cannot help but wonder: Why do different reasoning chains perform differently in reasoning? What components of the reasoning chains play a key role? Existing studies mainly focus on evaluating reasoning chains from a functional perspective, with little attention paid to their structural mechanisms. To address this gap, this work is the first to analyze and evaluate the quality of the reasoning chain from a structural perspective. We apply persistent homology from Topological Data Analysis (TDA) to map reasoning steps into semantic space, extract topological features, and analyze structural changes. These changes reveal semantic coherence, logical redundancy, and identify logical breaks and gaps. By calculating homology groups, we assess connectivity and redundancy at various scales, using barcode and persistence diagrams to quantify stability and consistency. Our results show that the topological structural complexity of reasoning chains correlates positively with accuracy. More complex chains identify correct answers sooner, while successful reasoning exhibits simpler topologies, reducing redundancy and cycles, enhancing efficiency and interpretability. This work provides a new perspective on reasoning chain quality assessment and offers guidance for future optimization.

大模型推理拓扑分析可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。