通过增强思维向量方差,提升大模型复杂推理能力。
LTA-thinker: Latent Thought-Augmented Training Framework for Large Language Models on Complex Reasoning
- 构建可学习先验的思维向量生成架构,扩大分布方差。
- 引入语义对齐与推理聚焦双损失,提升信息效率。
- 适合追求推理性能上限和高效训练的研究者。
大语言模型在复杂推理任务中可通过测试时缩放(TTS)动态优化,以缓解过度思考问题。现有方法如Coconut、SoftCoT及其变体在连续潜在空间推理中表现良好,但核心瓶颈仍在于高质量潜在思维向量的有效生成与利用。基于SoftCoT++理论——生成潜在思维分布的更大方差更逼近理想真实分布——我们提出潜思增强训练框架LTA-Thinker,从两方面提升分布方差并增强推理性能。首先,LTA-Thinker设计基于可学习先验的潜在思维生成架构,旨在增加生成思维向量的分布方差,简化结构并提升性能上限。其次,引入基于分布的方向优化范式,联合约束分布局部性与尺度,通过多目标协同训练策略,结合标准监督微调(SFT)损失与两项新损失:语义对齐损失(利用KL散度确保思维与问题语义高度相关)、推理聚焦损失(采用对比学习机制引导模型关注关键推理步骤)。实验表明,LTA-Thinker在多个基线中达到最先进(SOTA)性能,展现更高性能上限与更优缩放效应。
原文摘要 · Abstract (English)
Complex Reasoning in Large Language Models can be dynamically optimized using Test-Time Scaling (TTS) to mitigate Overthinking. Methods such as Coconut, SoftCoT and its variant are effective in continuous latent space inference, the core bottleneck still lies in the efficient generation and utilization of high-quality Latent Thought. Drawing from the theory of SoftCoT++ that a larger variance in the generated Latent Thought distribution more closely approximates the golden truth distribution, we propose a Latent Thought-Augmented Training Framework--LTA-Thinker, which improves distributional variance and enhances reasoning performance from two perspectives. First, LTA-Thinker constructs a Latent Thought generation architecture based on a learnable prior. This architecture aims to increase the variance distribution of generated Latent Thought Vectors in order to simplify the overall structure and raise the performance ceiling. Second, LTA-Thinker introduces a distribution-based directional optimization paradigm that jointly constrains both distribution locality and distribution scale. This mechanism improves information efficiency and computational cost through a multi-objective co-training strategy, which combines standard Supervised Fine-Tuning (SFT) loss with two novel losses: Semantic Alignment Loss, which utilizes KL divergence to ensure that the Latent Thought is highly relevant to the semantics of the question; Reasoning Focus Loss, which utilizes a contrastive learning mechanism to guide the model to focus on the most critical reasoning steps. Experiments show that LTA-thinker achieves state-of-the-art (SOTA) performance among various baselines and demonstrates a higher performance ceiling and better scaling effects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。