评测五种语音合成模型对数学表达式的可懂度,发现效果远不如人工朗读。
Intelligibility of Text-to-Speech Systems for Mathematical Expressions
- 用大语言模型将LaTeX转为发音,测试五种语音合成模型的输出质量。
- 多数数学表达式类别下,语音合成结果正确率远低于人工朗读。
- 不同模型和表达式类型间差异显著,需针对性优化语音合成系统。
现有研究对高级文本转语音(TTS)模型在数学表达式(MX)输入下的表现评估有限。本文设计实验,通过听觉测试与转写任务,评估五种TTS模型在多种数学表达式类别上的质量与可懂度。由于TTS模型无法直接处理LaTeX,我们采用两个大型语言模型(LLMs)将LaTeX数学表达式转换为英文发音作为输入。通过用户评分的平均意见分(MOS)以及三项转写正确性指标量化可懂度,并与人类专家朗读进行对比。结果表明,TTS模型生成的数学表达式语音并不必然可懂,不同模型及表达式类别间的可懂度差距明显。大多数类别中,TTS性能显著劣于人工朗读。大语言模型的选择影响较小。这表明亟需改进针对数学表达式的语音合成技术。
原文摘要 · Abstract (English)
There has been limited evaluation of advanced Text-to-Speech (TTS) models with Mathematical eXpressions (MX) as inputs. In this work, we design experiments to evaluate quality and intelligibility of five TTS models through listening and transcribing tests for various categories of MX. We use two Large Language Models (LLMs) to generate English pronunciation from LaTeX MX as TTS models cannot process LaTeX directly. We use Mean Opinion Score from user ratings and quantify intelligibility through transcription correctness using three metrics. We also compare listener preference of TTS outputs with respect to human expert rendition of same MX. Results establish that output of TTS models for MX is not necessarily intelligible, the gap in intelligibility varies across TTS models and MX category. For most categories, performance of TTS models is significantly worse than that of expert rendition. The effect of choice of LLM is limited. This establishes the need to improve TTS models for MX.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。