用拓扑分析方法自动评估大模型推理质量,更准更快。
The Shape of Reasoning: Topological Analysis of Reasoning Traces in Large Language Models
- 用拓扑数据分析推理路径的几何结构,而非简单看连接关系。
- 拓扑特征预测推理质量的效果远超传统图指标。
- 结果稳定可用,适合用于强化学习优化推理过程。
评估大语言模型推理轨迹的质量仍缺乏系统性方法,现有手段依赖专家评分、人工标注和耗时的成对判断,效率低且不可靠。自动化方法多基于图结构代理,仅衡量连通性,却无法解释高质量推理的本质;这类抽象过于简化,难以捕捉复杂推理过程的内在特征。本文提出一种基于拓扑数据分析(TDA)的评估框架,能捕捉推理轨迹的几何特性,实现标签高效、自动化的质量评估。实证研究表明,拓扑特征对推理质量的预测能力显著优于标准图度量,表明有效推理更应由高维几何结构刻画,而非单纯的关系图。进一步发现,一组精简且稳定的拓扑特征可可靠指示轨迹质量,为未来强化学习算法提供实用信号。
原文摘要 · Abstract (English)
Evaluating the quality of reasoning traces from large language models remains understudied, labor-intensive, and unreliable: current practice relies on expert rubrics, manual annotation, and slow pairwise judgments. Automated efforts are dominated by graph-based proxies that quantify structural connectivity but do not clarify what constitutes high-quality reasoning; such abstractions can be overly simplistic for inherently complex processes. We introduce a topological data analysis (TDA)-based evaluation framework that captures the geometry of reasoning traces and enables label-efficient, automated assessment. In our empirical study, topological features yield substantially higher predictive power for assessing reasoning quality than standard graph metrics, suggesting that effective reasoning is better captured by higher-dimensional geometric structures rather than purely relational graphs. We further show that a compact, stable set of topological features reliably indicates trace quality, offering a practical signal for future reinforcement learning algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。