小模型解数学题难?用提示协作来突破。
HintMR: Eliciting Stronger Mathematical Reasoning in Small Language Models

- 让小模型分步解题,每步由另一个小模型提供上下文提示。
- 在多个数学基准上,准确率显著高于传统提示方法。
- 适合想提升小模型推理能力的研究者或应用开发者。
小型语言模型(SLMs)因难以维持长链条中间步骤且易受早期错误影响,常在复杂数学推理中表现不佳。本文提出一种提示辅助推理框架,将解题过程分解为连续步骤,并由一个通过蒸馏训练的独立小模型生成上下文感知提示。该提示生成模型虽无法独立解题,但与推理模型协作,形成双模型协同系统。提示基于问题陈述和累积推理历史动态生成,提供分步、局部指导而不暴露完整解法,减少错误传播并帮助模型聚焦于可管理的子问题。跨多种数学基准和模型的实验表明,提示辅助显著提升小模型的推理准确率,相比标准提示有明显改进,同时保持模型效率。结果表明,通过提示生成与推理的小模型间结构化协作,是一种高效且轻量的增强数学推理机制。
原文摘要 · Abstract (English)
Small language models (SLMs) often struggle with complex mathematical reasoning due to limited capacity to maintain long chains of intermediate steps and to recover from early errors. We address this challenge by introducing a hint-assisted reasoning framework that incrementally guides SLMs through multi-step mathematical problem solving. Our approach decomposes solutions into sequential reasoning steps and provides context-aware hints, where hints are generated by a separate SLM trained via distillation from a strong large language model. While the hint-generating SLM alone is not capable of solving the problems, its collaboration with a reasoning SLM enables effective guidance, forming a cooperative two-model system for reasoning. Each hint is generated conditionally on the problem statement and the accumulated reasoning history, providing stepwise, localized guidance without revealing full solutions. This reduces error propagation and allows the reasoning model to focus on manageable subproblems. Experiments across diverse mathematical benchmarks and models demonstrate that hint assistance consistently improves reasoning accuracy for SLMs, yielding substantial gains over standard prompting while preserving model efficiency. These results highlight that structured collaboration between SLMs-via hint generation and reasoning-offers an effective and lightweight mechanism for enhancing mathematical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。