arXiv:2606.08974cs.AI2026-06

让大模型生成更多样化的推理路径,提升解题能力。

Diverse Thinking Schemata Elicit Better Reasoning in Large Language Models

论文配图:Diverse Thinking Schemata Elicit Better Reasoning in Large Language Models
图 1 · 摘自论文原文
  • 引入思维模式多样性概念,通过强化学习鼓励模型产生不同推理路径。
  • 在多个数学推理数据集上,准确率显著优于传统方法。
  • 适合研究大模型推理机制或想提升模型纠错能力的读者。

大型推理模型(LRMs)因其能生成长链条推理而受到关注,可解决复杂数学问题。本文聚焦推理过程中两个未被充分探索的方面:推理转移(步骤间的转换方式)和答案候选(模型产生的不同解题路径),统称为思维模式。我们发现思维模式多样性与模型性能正相关,由此提出多样化思维模式策略优化(DiScO)框架:首先赋予模型对思维模式的认知,再通过强化学习促进多样性,并在推理时进一步引导多样思考。在多个数学推理基准测试中,DiScO持续优于标准组相对策略优化方法。此外,人工标注分析显示,DiScO显著增强了模型从初始错误尝试中恢复的能力。结果表明,思维模式多样性至关重要,沿多样性方向扩展是未来重要研究方向。

原文摘要 · Abstract (English)

Large reasoning models (LRMs) have attracted increasing attention for their ability to solve complex mathematical problems by generating extended reasoning chains. In this work, we focus on two critical yet underexplored aspects of the reasoning process: reasoning transitions capturing the distinct transitions between reasoning steps and answer candidates reflecting the variety of solution paths produced by the model. We collectively define these two aspects as thinking schemata. We observe a correlation between the diversity of thinking schemata and model performance, which motivates us to enhance diversity as a means to further improve reasoning potential. To this end, we propose Diverse Schemata Policy Optimization (DiScO), a framework that first endows the model with schemata awareness, then encourages diversity through reinforcement learning, and further promotes diverse reasoning at inference time. Experiments on multiple mathematical reasoning benchmarks demonstrate that DiScO consistently outperforms standard group relative policy optimization. Beyond accuracy, human-annotated analyses show that DiScO substantially improves the model's ability to recover from erroneous initial attempts. Overall, our work suggests the important role that diversity of the thinking schemata plays and points to scaling along the diversity dimension as a promising research direction.

推理增强思维模式强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。