arXiv:2506.16174cs.LG2025-06

测试AI在转录芬兰说唱时的幻觉水平,对比两种语音识别模型表现。

Hallucination Level of Artificial Intelligence Whisperer: Case Speech Recognizing Pantterinousut Rap Song

  • 用真实芬兰说唱歌词作为参考,对比Faster Whisper与YouTube语音识别
  • 发现两者在合成音乐背景下的误识别率均较高,尤其对押韵和方言词
  • 适合关注语音识别鲁棒性及多语言生成评估的研究者

所有语言都有其独特性,其中芬兰语被认为是一种复杂语言。当艺术家使用语言时,发音与含义更难理解。本文将AI置于一项有趣而具有挑战性的任务中:将一首由作者弟弟Mc Timo创作的芬兰说唱歌曲转录为文字。我们对比Faster Whisper算法与YouTube内部语音转文本功能的表现。参考标准为原始芬兰说唱歌词,该歌曲由歌手在Syntikka Janne的合成音乐伴奏下演唱。通过比较AI输出与原始歌词之间的差异,评估其幻觉水平与误听程度。误差度量虽非正式,但在本案例中仍具有效性。

原文摘要 · Abstract (English)

All languages are peculiar. Some of them are considered more challenging to understand than others. The Finnish Language is known to be a complex language. Also, when languages are used by artists, the pronunciation and meaning might be more tricky to understand. Therefore, we are putting AI to a fun, yet challenging trial: translating a Finnish rap song to text. We will compare the Faster Whisperer algorithm and YouTube's internal speech-to-text functionality. The reference truth will be Finnish rap lyrics, which the main author's little brother, Mc Timo, has written. Transcribing the lyrics will be challenging because the artist raps over synth music player by Syntikka Janne. The hallucination level and mishearing of AI speech-to-text extractions will be measured by comparing errors made against the original Finnish lyrics. The error function is informal but still works for our case.

语音识别芬兰语幻觉检测说唱转录

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。