发现长思维链有类分子结构,可提升大模型推理能力
The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning
- 将长思维链比作类分子结构,由三种交互方式构成
- 只有促进熵快速收敛的结构才能稳定学习,否则训练受阻
- 提出Mole-Syn方法,显著提升多个基准上的推理性能
大型语言模型在模仿人类或非长思维链模型时,难以学会有效的长思维链(Long CoT)推理。为理解这一现象,我们提出有效且可学习的长思维链轨迹具有统一视角下的稳定类分子结构,由三种相互作用构成:深度推理(类共价键)、自我反思(类氢键)和自我探索(类范德华力)。对提炼出的思维链轨迹分析显示,这些结构源于长思维链微调,而非关键词模仿。我们引入有效语义异构体,发现仅促进快速熵收敛的连接支持稳定长思维链学习,而结构竞争会损害训练。基于此,我们提出Mole-Syn,一种分布转移图方法,指导生成有效长思维链结构,在多个基准上提升了性能与强化学习稳定性。
原文摘要 · Abstract (English)
Large language models (LLMs) often fail to learn effective long chain-of-thought (Long CoT) reasoning from human or non-Long-CoT LLMs imitation. To understand this, we propose that effective and learnable Long CoT trajectories feature stable molecular-like structures in unified view, which are formed by three interaction types: Deep-Reasoning (covalent-like), Self-Reflection (hydrogen-bond-like), and Self-Exploration (van der Waals-like). Analysis of distilled trajectories reveals these structures emerge from Long CoT fine-tuning, not keyword imitation. We introduce Effective Semantic Isomers and show that only bonds promoting fast entropy convergence support stable Long CoT learning, while structural competition impairs training. Drawing on these findings, we present Mole-Syn, a distribution-transfer-graph method that guides synthesis of effective Long CoT structures, boosting performance and RL stability across benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。