arXiv:2509.23219cs.LG2025-09被引 3

用强化学习让小模型学会无线通信数学,准确率接近大模型。

WirelessMathLM: Teaching Mathematical Reasoning for LLMs in Wireless Communications with Reinforcement Learning

  • 通过可验证正确性的奖励机制,无需人工标注训练小模型。
  • 70亿参数模型在4027题基准上达39.5%准确率,接近GPT-4o。
  • 训练后通用数学能力提升,零训练跨领域增8.4分。

大型语言模型在通用数学推理上表现优异,但在专业领域的技术数学中却表现灾难性失败。在无线通信领域,问题需精确处理信息论边界、优化约束和信号处理公式,即使最先进的模型也难以达到良好性能。本文提出WirelessMathLM,证明通过特定领域强化学习与可验证奖励,小型模型(0.5B–7B参数)可达到甚至超越更大模型的表现。关键洞察在于无线数学问题具有可验证正确性的独特属性,使强化学习无需人类反馈即可有效进行。我们构建了WirelessMathBench-XL,涵盖来自970篇论文的4,027道题目。采用组相对策略优化(GRPO)与二值验证奖励,从基础检查点直接训练模型,无需监督预热。我们的7B模型在WirelessMathBench-XL上达到39.5%准确率,接近GPT-4o(40.4%),而参数量仅为DeepSeek-R1(671B,57.4%)的约1/100。令人惊讶的是,GRPO训练使所有模型规模性能几乎翻倍(0.5B +11%,3B +103%,7B +81%),并展现出正向迁移:在MATH、Minerva-Math、OlympiadBench、AMC和AIME等通用数学基准上平均提升+8.4分,且未在此类任务上进行任何训练。

原文摘要 · Abstract (English)

Large language models (LLMs) excel at general mathematical reasoning but fail catastrophically on specialized technical mathematics. In wireless communications, where problems require precise manipulation of information-theoretic bounds, optimization constraints, and signal processing formulations, even state-of-the-art models struggle to achieve competent performance. We present WirelessMathLM, demonstrating that compact models (0.5B-7B parameters) can match or exceed much larger models through domain-specific reinforcement learning with verifiable rewards. Our key insight is that wireless mathematics problems possess a unique property--verifiable correctness--that enables effective reinforcement learning without human feedback. We construct WirelessMathBench-XL, a comprehensive benchmark of 4,027 problems from 970 papers. Using Group Relative Policy Optimization (GRPO) with binary verification rewards, we train models directly from base checkpoints without supervised warm-start. Our 7B model achieves 39.5% accuracy on WirelessMathBench-XL, approaching GPT-4o (40.4%) while using about 100 times fewer parameters than DeepSeek-R1 (671B, 57.4%). Remarkably, GRPO training nearly doubles performance across all model scales (0.5B +11%, 3B +103%, 7B +81%), with positive transfer to general mathematics benchmarks--our models gain +8.4 points on average across MATH, Minerva-Math, OlympiadBench, AMC, and AIME without any training on these tasks.

数学推理强化学习无线通信小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。