arXiv:2504.05702cs.CL2025-04被引 3

用诗歌朗读音频评估语音转文字系统,发现Whisper开源最佳但需防幻觉。

Evaluating Speech-to-Text Systems with PennSound

  • 用彭声库10小时多变语音数据测试多个系统,构建可靠基准
  • Rev.ai错词率最低,Whisper在开源中表现最好且可避免幻觉
  • 适合语音识别研究者、开发者及需高鲁棒性系统的用户

从全球最大的诗歌朗读与讨论在线资源PennSound中随机抽取近10小时语音,作为基准评估多个商业与开源语音转文字系统。该数据集涵盖多样录音条件与语音风格,代表众多未转录音频集合。参考转写由训练标注员完成,系统转写来自AWS、Azure、Google、IBM、NeMo、Rev.ai、Whisper和Whisper.cpp。基于词错率(WER),Rev.ai表现最优;Whisper在开源系统中领先,前提是避免幻觉。AWS在三个系统中拥有最佳说话人分离错误率(DER)。尽管WER与DER差异微小,但不同权衡可能影响实际选择。研究还分析了Whisper中的幻觉问题,提醒用户注意运行时选项,以及速度与准确率的取舍是否可接受。

原文摘要 · Abstract (English)

A random sample of nearly 10 hours of speech from PennSound, the world's largest online collection of poetry readings and discussions, was used as a benchmark to evaluate several commercial and open-source speech-to-text systems. PennSound's wide variation in recording conditions and speech styles makes it a good representative for many other untranscribed audio collections. Reference transcripts were created by trained annotators, and system transcripts were produced from AWS, Azure, Google, IBM, NeMo, Rev.ai, Whisper, and Whisper.cpp. Based on word error rate, Rev.ai was the top performer, and Whisper was the top open source performer (as long as hallucinations were avoided). AWS had the best diarization error rates among three systems. However, WER and DER differences were slim, and various tradeoffs may motivate choosing different systems for different end users. We also examine the issue of hallucinations in Whisper. Users of Whisper should be cautioned to be aware of runtime options, and whether the speed vs accuracy trade off is acceptable.

语音识别评测基准Whisper幻觉检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。