arXiv:2508.03542cs.CV2025-08被引 2

首个开源数学语音转LaTeX数据集,支持中英文方程与句子转换。

Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences

  • 构建66000+条多语言数学语音数据,涵盖方程与句子
  • 在新基准上方程转LaTeX错误率降至27%,领先36个百分点
  • 首次建立数学句子识别基准,适合教育与科研场景

将口语化数学表达转化为严格结构化的符号表示是一项挑战,需解决发音歧义问题。尽管自动语音识别(ASR)和语言模型(LM)已取得进展,但口语数学转LaTeX仍研究不足。本文提出首个完全开源的大规模数据集,包含超过66,000条英语与俄语的数学方程及句子音频样本,覆盖多样科学领域。基于ASR后纠错、少量样本提示及音频语言模型,我们在MathSpeech基准上实现28%字符错误率(CER),接近现有模型的30%;在新提出的S2L-equations基准上,错误率降至27%,较MathSpeech模型提升超36个百分点(64%)。我们还建立了首个数学句子识别基准(S2L-sentences),方程转LaTeX CER达40%。该工作为多模态AI中的数学内容识别奠定基础。

原文摘要 · Abstract (English)

Conversion of spoken mathematical expressions is a challenging task that involves transcribing speech into a strictly structured symbolic representation while addressing the ambiguity inherent in the pronunciation of equations. Although significant progress has been achieved in automatic speech recognition (ASR) and language models (LM), the problem of converting spoken mathematics into LaTeX remains underexplored. This task directly applies to educational and research domains, such as lecture transcription or note creation. Based on ASR post-correction, prior work requires 2 transcriptions, focuses only on isolated equations, has a limited test set, and provides neither training data nor multilingual coverage. To address these issues, we present the first fully open-source large-scale dataset, comprising over 66,000 human-annotated audio samples of mathematical equations and sentences in English and Russian, drawn from diverse scientific domains. In addition to the ASR post-correction models and few-shot prompting, we apply audio language models, demonstrating comparable character error rate (CER) results on the MathSpeech benchmark (28% vs. 30%) for the equations conversion. In contrast, on the proposed S2L-equations benchmark, our models outperform the MathSpeech model by a substantial margin of more than 36 percentage points, even after accounting for LaTeX formatting artifacts (27% vs. 64%). We establish the first benchmark for mathematical sentence recognition (S2L-sentences) and achieve an equation CER of 40%. This work lays the groundwork for future advances in multimodal AI, with a particular focus on mathematical content recognition.

语音转公式数学识别开源数据多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。