arXiv:2505.14354cs.CLcs.LG2025-05ACL被引 12

评测大模型在无线通信数学建模中的推理能力,发现其在方程补全上表现很差。

WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications

  • 构建587道来自40篇顶刊的无线通信数学题,涵盖选择与方程补全
  • 顶尖模型仅38.05%平均准确率,完整方程补全成功率不足8%
  • 专为无线工程设计,适合研究大模型在专业领域推理的学者

大语言模型在众多任务中表现优异,但在无线通信等专业领域的复杂数学推理能力仍不清晰。本文提出WirelessMathBench,一个针对无线通信工程数学建模的新型评测基准。该基准包含587道精心筛选的问题,源自40篇顶级论文,涵盖从基础选择题到复杂方程补全(部分与完整)的多样化任务,均严格满足物理与量纲约束。对主流LLMs的大量实验表明,尽管多数模型在基础记忆任务中表现良好,但在部分或完整方程重建任务中性能显著下降,暴露出当前模型的根本局限。即使表现最佳的DeepSeek-R1,平均准确率也仅为38.05%,完整方程补全成功率仅为7.83%。我们公开发布WirelessMathBench及评估工具包,旨在推动更鲁棒、领域感知更强的LLMs在无线系统分析与更广泛工程应用中的发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved impressive results across a broad array of tasks, yet their capacity for complex, domain-specific mathematical reasoning-particularly in wireless communications-remains underexplored. In this work, we introduce WirelessMathBench, a novel benchmark specifically designed to evaluate LLMs on mathematical modeling challenges to wireless communications engineering. Our benchmark consists of 587 meticulously curated questions sourced from 40 state-of-the-art research papers, encompassing a diverse spectrum of tasks ranging from basic multiple-choice questions to complex equation completion tasks, including both partial and full completions, all of which rigorously adhere to physical and dimensional constraints. Through extensive experimentation with leading LLMs, we observe that while many models excel in basic recall tasks, their performance degrades significantly when reconstructing partially or fully obscured equations, exposing fundamental limitations in current LLMs. Even DeepSeek-R1, the best performer on our benchmark, achieves an average accuracy of only 38.05%, with a mere 7.83% success rate in full equation completion. By publicly releasing WirelessMathBench along with the evaluation toolkit, we aim to advance the development of more robust, domain-aware LLMs for wireless system analysis and broader engineering applications.

数学推理无线通信大模型评测方程补全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。