首个评估语音数学推理能力的基准,揭示语音模型在算术与符号理解上的短板。
Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems
- 构建语音数学问答基准,涵盖算术、多步推理等多样题型
- 语音模型在基础算术任务上表现仍差,对口语化表达理解弱
- 对LaTeX符号有偏倚,知识推理能力在语音输入下大幅下降
大语言模型(LLMs)和多模态大语言模型(MLLMs)在多项任务中展现出强大的推理能力,但其从语音输入进行数学推理的能力仍缺乏研究。以往语音研究多集中于事实理解或简单音频推理,难以反映数学问题求解所需的逻辑步骤推理。为此,我们提出语音数学问答(Spoken-MQA)基准,用于评估语音模型(包括级联模型和端到端语音LLM)的数学推理能力。该基准涵盖纯算术、单步与多步情境推理、知识导向推理等多种题型,所有题目均以清晰自然的口语形式呈现。实验发现:(1) 某些语音LLM在涉及基本算术的情境推理任务中表现良好,但在直接算术任务上仍存在困难;(2) 当前LLM对LaTeX书写的数学符号有显著偏好,难以理解口语化数学表达;(3) 语音输入下数学知识推理能力明显退化。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) and multimodal LLMs (MLLMs) have led to strong reasoning ability across a wide range of tasks. However, their ability to perform mathematical reasoning from spoken input remains underexplored. Prior studies on speech modality have mostly focused on factual speech understanding or simple audio reasoning tasks, providing limited insight into logical step-by-step reasoning, such as that required for mathematical problem solving. To address this gap, we introduce Spoken Math Question Answering (Spoken-MQA), a new benchmark designed to evaluate the mathematical reasoning capabilities of speech-based models, including both cascade models (ASR + LLMs) and end-to-end speech LLMs. Spoken-MQA covers a diverse set of math problems, including pure arithmetic, single-step and multi-step contextual reasoning, and knowledge-oriented reasoning problems, all presented in unambiguous natural spoken language. Through extensive experiments, we find that: (1) while some speech LLMs perform competitively on contextual reasoning tasks involving basic arithmetic, they still struggle with direct arithmetic problems; (2) current LLMs exhibit a strong bias toward symbolic mathematical expressions written in LaTex and have difficulty interpreting verbalized mathematical expressions; and (3) mathematical knowledge reasoning abilities are significantly degraded in current speech LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。