arXiv:2506.10674cs.AIcs.CL2025-06被引 4

首个面向通信领域数学题的评测基准,测试大模型实际解题能力。

TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving

  • 基于专家设计的题目生成500道通信数学题
  • 专精推理模型表现远超通用大模型
  • 适合评估大模型在工程场景中的数学能力

人工智能在通信领域的应用日益广泛,激发了对大语言模型(LLMs)处理特定领域、高数学强度任务能力的关注。尽管近期进展提升了LLMs在通用数学推理方面的能力,但在信号处理、网络优化和性能分析等专业领域,其有效性仍缺乏系统研究。为此,我们提出TeleMath,首个专门用于评估大语言模型在通信领域求解数学问题(含数值答案)的基准数据集。该数据集包含500个问答对,覆盖通信领域的广泛主题。本文介绍了从领域专家设计的种子问题出发的问答生成流程。对多种开源大模型的评估显示,专为数学或逻辑推理设计的最新模型在TeleMath上表现最佳;而即使参数量巨大,通用模型也常在此类任务中表现不佳。我们已公开数据集与评估代码,以促进结果复现并支持后续研究。

原文摘要 · Abstract (English)

The increasing adoption of artificial intelligence in telecommunications has raised interest in the capability of Large Language Models (LLMs) to address domain-specific, mathematically intensive tasks. Although recent advancements have improved the performance of LLMs in general mathematical reasoning, their effectiveness within specialized domains, such as signal processing, network optimization, and performance analysis, remains largely unexplored. To address this gap, we introduce TeleMath, the first benchmark dataset specifically designed to evaluate LLM performance in solving mathematical problems with numerical solutions in the telecommunications domain. Comprising 500 question-answer (QnA) pairs, TeleMath covers a wide spectrum of topics in the telecommunications field. This paper outlines the proposed QnAs generation pipeline, starting from a selected seed of problems crafted by Subject Matter Experts. The evaluation of a wide range of open-source LLMs reveals that best performance on TeleMath is achieved by recent models explicitly designed for mathematical or logical reasoning. In contrast, general-purpose models, even those with a large number of parameters, often struggle with these challenges. We have released the dataset and the evaluation code to ease result reproducibility and support future research.

大模型评测通信数学推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。