用可学习的思维向量让大模型数学推理更可控、更专注。
Controllable Mathematical Reasoning via Self-Optimizing Thought Vectors
- 引入可学习的思维向量动态调节大模型推理过程。
- 在GSM8K上达90.1%准确率,可控性得分0.42。
- 无需外部奖励标注,靠熵最小化实现专注推理。
我们提出一种新的可控制数学推理方法,利用基于熵最小化的自优化思维向量。该方法引入可学习的思维向量,动态调节大语言模型的内部推理过程。在GSM8K数据集上使用Gemma-2-9B模型,取得90.1%的准确率和0.42的可控性得分,证明熵激励能有效引导聚焦的推理模式,且无需外部奖励标注。分析显示不同控制条件下存在明确的思维向量聚类和稳定的低熵分布,验证了该框架在可控AI推理中的有效性。
原文摘要 · Abstract (English)
We present a novel approach for controllable mathematical reasoning that leverages self-optimizing thought vectors with entropy minimization. Our method introduces learnable thought vectors that dynamically modulate the internal reasoning process of large language models. Using Gemma-2-9B on GSM8K, we achieve 90.1% accuracy with a controllability score of 0.42, demonstrating that entropy-based rewards effectively guide focused reasoning patterns without requiring external reward annotations. Our analysis reveals distinct thought vector clusters and consistent low-entropy distributions across control conditions, validating our framework for controllable AI reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。