arXiv:2608.17301cs.AI2026-08

30亿参数模型经强化训练后,在信号数学题上准确率提升至39.12%。

SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning

论文配图:SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning
图 1 · 摘自论文原文
  • 用可验证奖励进行强化学习,结合领域特定思维链微调
  • 最佳模型准确率达39.12%,是基础模型的三倍以上
  • 适用于信号处理领域的数学推理研究者

通过监督式思维链微调与可验证奖励的强化学习,大语言模型的数学推理能力显著提升。然而其在信号处理问题中的应用仍较有限。本报告研究将Qwen2.5-3B-Base模型适配至无线通信领域高等级信号数学问题的方法,基于WirelessMATHBench-XL基准测试。考察两种训练范式:(i) 直接在WirelessMATHBench-XL上使用可验证奖励进行强化学习;(ii) 先在精炼的无线领域思维链语料上进行监督微调,再进行相同领域的强化学习。对比了Group Relative Policy Optimization (GRPO)、Group Sequence Policy Optimization (GSPO) 和 Geometric-Mean Policy Optimization (GMPO) 三种优化算法。旨在评估领域感知思维链微调是否能有效作为后续强化学习的初始化,并比较GSPO与GMPO在稳定性与准确性上是否优于GRPO。最佳模型整体准确率达到39.12%,较未经训练的Base模型(12.37%)提升超过三倍。

原文摘要 · Abstract (English)

Post-training with supervised chain-of-thought fine-tuning and reinforcement learning from verifiable rewards has substantially improved the mathematical reasoning capabilities of large language models (LLMs). However, their application to signal processing problems remains relatively under-explored. This report investigates reinforcement fine-tuning strategies for adapting Qwen2.5-3B-Base to graduate-level signal mathematical problems from WirelessMATHBench-XL, a comprehensive benchmark for mathematical reasoning in this domain. We examine two training paradigms: (i) direct reinforcement learning (RL) on WirelessMATHBench-XL with verifiable rewards; and (ii) supervised fine-tuning (SFT) on a distilled wireless-domain chain-of-thought corpus, followed by the same domain-specific RL stage. Across both paradigms, we benchmark Group Relative Policy Optimization (GRPO), Group Sequence Policy Optimization (GSPO), and Geometric-Mean Policy Optimization (GMPO). We aim to assess whether domain-aware CoT SFT serves as an effective initialization for subsequent RL, and whether GSPO or GMPO offer advantages in stability or accuracy over GRPO for signal reasoning tasks. Our best model achieves an overall accuracy of 39.12\%, representing a more than threefold improvement over the untrained Base model (12.37\%).

信号处理数学推理强化学习小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。